Migrate legacy HTTPConnection callers to the policy-aware Transport incrementally
维护者通常 1 天内回复
@AbhiPrasad 已经在做这个了。
开始于 2026年10月1日。
评估
这个 Issue 还没有评估数据。
描述
Problem
The Python SDK has two HTTP request implementations: the policy-aware braintrust.api._transport.Transport used by authentication and generated REST resources, and legacy HTTPConnection sessions used by remaining high-level SDK paths. Keeping both duplicates connection setup, timeout configuration, retry behavior, and lifecycle handling.
The legacy /logs3 writer illustrates the operational cost:
- It retries each batch independently, with three attempts by default and 1-second/2-second backoff, without honoring
Retry-After. - Batches run concurrently in a thread pool sized to
cpu_count(), without a shared throttling cooldown. - Terminal batch failures print to stderr and return normally unless
BRAINTRUST_SYNC_FLUSHis enabled. Callingbraintrust.flush()does not itself enable that mode. - The writer drains its queue before submitting batches, so exhausted requests can discard score updates while callers believe flushing succeeded.
The new transport already provides operation-specific retry policies, bounded retries, Retry-After parsing, and structured errors. However, RetryMode.LOG_INGESTION deliberately performs a single attempt, and the transport has no shared cooldown. Moving the writer to it does not, by itself, fix failed-batch retention or persistence reporting.
Proposed approach
Incrementally migrate callers directly to Transport, then remove HTTPConnection. Do not introduce a replacement compatibility wrapper around the legacy class as the migration architecture.
Reuse the existing policy-aware client in BraintrustState. Add small internal services/helpers for endpoints outside the generated REST API so routing, per-request authentication, and explicit operation policies stay centralized.
Use a common request engine and configuration; separate sessions may remain where credentials, concurrency, or signed storage requests require them. Signed storage requests must not inherit Braintrust authorization headers.
Optimize for log ingestion
Preserve RetryMode.LOG_INGESTION as a distinct mode and make efficient, reliable delivery to /logs3 a primary design requirement of the migration. Keep transport-backed ingestion attempts single-attempt by default, with the ingestion writer owning batch retries, scheduling, and retention throughout the migration and afterward.
- Keep ingestion-specific timeout and scheduling choices explicit rather than inheriting safe-read or generic write defaults.
- Preserve efficient batching by item count and payload bytes, serialize each batch once, and reuse the prepared payload across retries. Preserve the overflow-upload flow without unnecessary repeated uploads.
- Reuse persistent connections with connection-pool capacity appropriate to bounded ingestion concurrency. Allow an ingestion-specific pool/session when needed, while sharing the policy-aware request engine.
- Make in-flight ingestion concurrency explicit and bounded; do not rely solely on
cpu_count()as the concurrency policy. Ensure ingestion activity does not starve unrelated API calls. - Apply destination-scoped cooldowns to both retried and newly queued batches, and stagger recovery after cooldown so workers do not all resume in a burst.
- Preserve row identities and update ordering when replaying failed batches; verify the server's replay semantics rather than treating ingestion as an ordinary idempotent write.
- Measure healthy-path throughput, producer-side logging latency, allocation/serialization overhead, and flush latency against the current writer. Under throttling, measure rejected request volume, queue growth, and recovery to confirm that lower request pressure does not come from silently dropping batches.
Minimum PR plan: two SDK PRs
Use the existing coexistence of the policy-aware client and legacy connections to migrate in two steps. Include helpers, tests, benchmarks, documentation, and cleanup in the PR that needs them rather than opening standalone foundation or cleanup PRs.
PR 1: Optimize and harden /logs3 on the policy-aware transport
Deliver the log ingestion path as one coherent change:
- Add only the internal ingestion helpers and transport configuration needed by the writer, reusing existing authentication and routing.
- Migrate
/logs3, its version/payload-limit lookup, and its overflow URL/upload flow toTransport. Preserve the distinct, single-attemptLOG_INGESTIONmode; the writer owns retries, scheduling, and batch retention. - Implement destination-scoped
Retry-Aftercooldowns, bounded in-flight concurrency, staggered recovery, failed-batch retention/backpressure, preserved row ordering, and explicit flush failure reporting. - Preserve batching, prepared payload reuse, overflow upload reuse, and persistent connections; include regression tests and healthy/throttled performance measurements in this PR.
- Keep unrelated legacy callers working while using ingestion-specific pooling where needed.
Acceptance: ingestion uses the new engine, new and retried batches respect cooldowns, failed batches remain accounted for with visible capacity/permanent-failure handling, explicit flush reports unsuccessful delivery, and performance measurements show the effect on throughput and logging latency. Cover permanent 413s, long cooldowns, retry exhaustion, concurrent recovery, and relogin/adapter lifecycle.
PR 2: Migrate remaining callers and remove the legacy implementation
Build on PR 1's helpers/configuration where applicable:
- Migrate all remaining reads/BTQL, invocation/streaming, sandbox and framework publishing, generic attachments, CLI, devserver, and BTX direct requests.
- Classify endpoint retry safety explicitly and preserve special response/error handling, credential isolation, timeouts, application polling, and adapter ownership.
- Remove SDK-owned
HTTPConnectionrequest paths and obsolete connection setup, fields, factories, adapter retries, and tests. Move token sanitation out of the legacy class so policy-aware auth/client code no longer imports it. - Include documentation and any required compatibility/deprecation treatment for supported external extension points; retain such surfaces only as explicitly documented compatibility exceptions.
- Run the affected caller tests, broad core coverage, and static checks with this PR.
Acceptance: no SDK-owned request uses HTTPConnection, compatibility exceptions are documented, and the legacy class can be deleted once external compatibility requirements permit.
PR 2 depends on PR 1. A one-PR implementation is technically possible, but two keeps ingestion scheduling/retention changes independently reviewable from the broad caller migration. Only split further if implementation reveals an actual compatibility blocker or the retention changes cannot be reviewed together with ingestion scheduling. Service-side score retention and OpenTelemetry exporter work remain outside these two SDK PRs.
Migration scope checklist
- Establish internal helpers and consistent transport configuration, including
BRAINTRUST_HTTP_TIMEOUT, session ownership, adapter replacement, and login/credential lifecycle. - Migrate reads and logical POST reads, including BTQL and remaining experiment-fetch paths.
- Migrate function invocation and sandbox helpers, preserving streaming behavior and specialized invocation errors.
- Migrate attachment metadata, uploads/downloads, and
/logs3overflow uploads, preserving signed headers and accounting for replayable versus multipart/file-like bodies. - Migrate CLI, devserver, framework function publishing, and direct
requestscalls in BTX helpers. Keep application-level polling distinct from HTTP retries. - Migrate the
/logs3writer to transport-backed individual attempts usingLOG_INGESTION, with the writer retaining ownership of batch retries, scheduling, and retention. - Remove unused legacy connection fields/factories,
make_long_lived()behavior, andRetryRequestExceptionsAdapterwhere supported API compatibility permits. - Remove
HTTPConnectionafter all SDK-owned callers have migrated; document/deprecate supported external connection extension points as needed.
Keep legacy connections only for callers that have not yet migrated. This checklist describes scope, not one PR per caller group; use the two-PR plan above, with focused commits and appropriate regression coverage within each PR.
Behavioral requirements
- Classify retry safety explicitly: safe reads, verified idempotent writes, non-retryable writes/executions, and log ingestion. Do not infer safety from the HTTP method alone or automatically replay function executions.
- Preserve response/error contracts, including the base-experiment 400 sentinel and
BraintrustInvokeErrorhandling for invocation 500s; the new transport raises HTTP errors before callers can inspect raw failure responses. - Preserve streaming behavior without restarting a partially consumed response.
- Preserve custom
set_http_adapter()support and caller-owned adapter/session lifecycle. Keep one retry owner per request to avoid multiplying adapter, transport, and writer retries. - Preserve credential isolation and organization routing across login changes and loader caches.
Separate /logs3 reliability milestone
Track these behavioral requirements separately from mechanical caller migration, and implement them together with the ingestion migration in PR 1:
- Honor
Retry-Afterand introduce a thread-safe cooldown shared by batches for the same ingestion destination. New batches must respect an active cooldown too. - Keep bounded retry budgets and avoid multiplying transport retries underneath writer retries.
- Retain/requeue failed batches with bounded backpressure and preserved update ordering; handle permanent failures explicitly instead of retrying indefinitely.
- Make queue-capacity exhaustion visible. The current
LogQueueuses a bounded deque that can evict older records silently, so blindly putting failed records back onto it is insufficient for reliable retention. - Make explicit
flush()report unsuccessful delivery and preserve useful error details, without requiringBRAINTRUST_SYNC_FLUSH. - Cover throttling across concurrent batches, cooldowns longer than the retry budget, retry exhaustion, and existing terminal 413 behavior.
Service-side retention/requeueing of scores when persistence fails is a separate service change. Durable recovery across SDK process crashes is also a separate scope decision. OpenTelemetry's exporter owns its transport and needs a separate assessment if it is to be included; provider SDK networking is outside this migration.
Validation and completion criteria
- Add failing regressions before behavior fixes, using the existing local HTTP test server for transport behavior and existing cassette-backed coverage where appropriate.
- Validate each migrated caller group narrowly first, then run
test_core, the relevant nox sessions, and static checks frompy/usingmiseand the authoritative nox/CI configuration. - Cover special errors, streaming, uploads, custom adapters, timeouts, credential changes, and connection cleanup.
- Preserve the distinct
LOG_INGESTIONmode and validate ingestion performance with the repo's SDK benchmarking workflow, alongside concurrent throttling and recovery tests. - No SDK-owned legacy
HTTPConnectionrequest path remains after migration, and any retained connection API is explicitly documented for compatibility. - Do not claim the throttling/data-loss issue is resolved until the separate ingestion reliability requirements are met.
Source references
Inspected at checkout commit 6537ec61d7809e9c06e8ba9d1955ca6b3eb53cac:
- 主要语言
- Python
- 星标
- 21
- 派生
- 23
- 平均合并
- 20 小时 8 分钟
- 30 天内合并 PR
- 94
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
braintrustdata/braintrust-sdk-python 的其他 Issue
-
feature python
难度 2/5 1-3 小时 新手友好度 76/100
braintrustdata/braintrust-sdk-python#868 ·
维护者通常 1 天内回复
-
feature integration: mistral python
难度 2/5 1-3 小时 新手友好度 75/100
braintrustdata/braintrust-sdk-python#867 ·
维护者通常 1 天内回复
-
[bot] Google GenAI: streaming responses drop `url_context_metadata` that non-streaming responses preserve可能已有人在做 @Kayvan-Zahiri 于 4 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 86/100
braintrustdata/braintrust-sdk-python#774 ·
维护者通常 1 天内回复
-
OpenAI: trace computer-use Agents session items and capture Responses `access_programs`可能已有人在做 @AbhiPrasad 于 1 天前认领。 未关闭feature integration: openai python
braintrustdata/braintrust-sdk-python#869 · 已指派 1 人 ·
维护者通常 1 天内回复
-
feature python
难度 3/5 1-2 天 新手友好度 72/100
braintrustdata/braintrust-sdk-python#866 ·
维护者通常 1 天内回复
查看 braintrustdata/braintrust-sdk-python 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 85/100
-
难度 1/5 1 小时以内 新手友好度 75/100
-
难度 1/5 1 小时以内 新手友好度 85/100
-
难度 1/5 1 小时以内 新手友好度 85/100
data-umbrella/du-event-board#225 ·
-
难度 2/5 1-3 小时 新手友好度 75/100