Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Migrate legacy HTTPConnection callers to the policy-aware Transport incrementally

未关闭
#839 0 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 1 天内回复

@AbhiPrasad 已经在做这个了。

开始于 2026年10月1日。

评估

这个 Issue 还没有评估数据。

描述

python

Problem

The Python SDK has two HTTP request implementations: the policy-aware braintrust.api._transport.Transport used by authentication and generated REST resources, and legacy HTTPConnection sessions used by remaining high-level SDK paths. Keeping both duplicates connection setup, timeout configuration, retry behavior, and lifecycle handling.

The legacy /logs3 writer illustrates the operational cost:

  • It retries each batch independently, with three attempts by default and 1-second/2-second backoff, without honoring Retry-After.
  • Batches run concurrently in a thread pool sized to cpu_count(), without a shared throttling cooldown.
  • Terminal batch failures print to stderr and return normally unless BRAINTRUST_SYNC_FLUSH is enabled. Calling braintrust.flush() does not itself enable that mode.
  • The writer drains its queue before submitting batches, so exhausted requests can discard score updates while callers believe flushing succeeded.

The new transport already provides operation-specific retry policies, bounded retries, Retry-After parsing, and structured errors. However, RetryMode.LOG_INGESTION deliberately performs a single attempt, and the transport has no shared cooldown. Moving the writer to it does not, by itself, fix failed-batch retention or persistence reporting.

Proposed approach

Incrementally migrate callers directly to Transport, then remove HTTPConnection. Do not introduce a replacement compatibility wrapper around the legacy class as the migration architecture.

Reuse the existing policy-aware client in BraintrustState. Add small internal services/helpers for endpoints outside the generated REST API so routing, per-request authentication, and explicit operation policies stay centralized.

Use a common request engine and configuration; separate sessions may remain where credentials, concurrency, or signed storage requests require them. Signed storage requests must not inherit Braintrust authorization headers.

Optimize for log ingestion

Preserve RetryMode.LOG_INGESTION as a distinct mode and make efficient, reliable delivery to /logs3 a primary design requirement of the migration. Keep transport-backed ingestion attempts single-attempt by default, with the ingestion writer owning batch retries, scheduling, and retention throughout the migration and afterward.

  • Keep ingestion-specific timeout and scheduling choices explicit rather than inheriting safe-read or generic write defaults.
  • Preserve efficient batching by item count and payload bytes, serialize each batch once, and reuse the prepared payload across retries. Preserve the overflow-upload flow without unnecessary repeated uploads.
  • Reuse persistent connections with connection-pool capacity appropriate to bounded ingestion concurrency. Allow an ingestion-specific pool/session when needed, while sharing the policy-aware request engine.
  • Make in-flight ingestion concurrency explicit and bounded; do not rely solely on cpu_count() as the concurrency policy. Ensure ingestion activity does not starve unrelated API calls.
  • Apply destination-scoped cooldowns to both retried and newly queued batches, and stagger recovery after cooldown so workers do not all resume in a burst.
  • Preserve row identities and update ordering when replaying failed batches; verify the server's replay semantics rather than treating ingestion as an ordinary idempotent write.
  • Measure healthy-path throughput, producer-side logging latency, allocation/serialization overhead, and flush latency against the current writer. Under throttling, measure rejected request volume, queue growth, and recovery to confirm that lower request pressure does not come from silently dropping batches.

Minimum PR plan: two SDK PRs

Use the existing coexistence of the policy-aware client and legacy connections to migrate in two steps. Include helpers, tests, benchmarks, documentation, and cleanup in the PR that needs them rather than opening standalone foundation or cleanup PRs.

PR 1: Optimize and harden /logs3 on the policy-aware transport

Deliver the log ingestion path as one coherent change:

  • Add only the internal ingestion helpers and transport configuration needed by the writer, reusing existing authentication and routing.
  • Migrate /logs3, its version/payload-limit lookup, and its overflow URL/upload flow to Transport. Preserve the distinct, single-attempt LOG_INGESTION mode; the writer owns retries, scheduling, and batch retention.
  • Implement destination-scoped Retry-After cooldowns, bounded in-flight concurrency, staggered recovery, failed-batch retention/backpressure, preserved row ordering, and explicit flush failure reporting.
  • Preserve batching, prepared payload reuse, overflow upload reuse, and persistent connections; include regression tests and healthy/throttled performance measurements in this PR.
  • Keep unrelated legacy callers working while using ingestion-specific pooling where needed.

Acceptance: ingestion uses the new engine, new and retried batches respect cooldowns, failed batches remain accounted for with visible capacity/permanent-failure handling, explicit flush reports unsuccessful delivery, and performance measurements show the effect on throughput and logging latency. Cover permanent 413s, long cooldowns, retry exhaustion, concurrent recovery, and relogin/adapter lifecycle.

PR 2: Migrate remaining callers and remove the legacy implementation

Build on PR 1's helpers/configuration where applicable:

  • Migrate all remaining reads/BTQL, invocation/streaming, sandbox and framework publishing, generic attachments, CLI, devserver, and BTX direct requests.
  • Classify endpoint retry safety explicitly and preserve special response/error handling, credential isolation, timeouts, application polling, and adapter ownership.
  • Remove SDK-owned HTTPConnection request paths and obsolete connection setup, fields, factories, adapter retries, and tests. Move token sanitation out of the legacy class so policy-aware auth/client code no longer imports it.
  • Include documentation and any required compatibility/deprecation treatment for supported external extension points; retain such surfaces only as explicitly documented compatibility exceptions.
  • Run the affected caller tests, broad core coverage, and static checks with this PR.

Acceptance: no SDK-owned request uses HTTPConnection, compatibility exceptions are documented, and the legacy class can be deleted once external compatibility requirements permit.

PR 2 depends on PR 1. A one-PR implementation is technically possible, but two keeps ingestion scheduling/retention changes independently reviewable from the broad caller migration. Only split further if implementation reveals an actual compatibility blocker or the retention changes cannot be reviewed together with ingestion scheduling. Service-side score retention and OpenTelemetry exporter work remain outside these two SDK PRs.

Migration scope checklist

  • Establish internal helpers and consistent transport configuration, including BRAINTRUST_HTTP_TIMEOUT, session ownership, adapter replacement, and login/credential lifecycle.
  • Migrate reads and logical POST reads, including BTQL and remaining experiment-fetch paths.
  • Migrate function invocation and sandbox helpers, preserving streaming behavior and specialized invocation errors.
  • Migrate attachment metadata, uploads/downloads, and /logs3 overflow uploads, preserving signed headers and accounting for replayable versus multipart/file-like bodies.
  • Migrate CLI, devserver, framework function publishing, and direct requests calls in BTX helpers. Keep application-level polling distinct from HTTP retries.
  • Migrate the /logs3 writer to transport-backed individual attempts using LOG_INGESTION, with the writer retaining ownership of batch retries, scheduling, and retention.
  • Remove unused legacy connection fields/factories, make_long_lived() behavior, and RetryRequestExceptionsAdapter where supported API compatibility permits.
  • Remove HTTPConnection after all SDK-owned callers have migrated; document/deprecate supported external connection extension points as needed.

Keep legacy connections only for callers that have not yet migrated. This checklist describes scope, not one PR per caller group; use the two-PR plan above, with focused commits and appropriate regression coverage within each PR.

Behavioral requirements

  • Classify retry safety explicitly: safe reads, verified idempotent writes, non-retryable writes/executions, and log ingestion. Do not infer safety from the HTTP method alone or automatically replay function executions.
  • Preserve response/error contracts, including the base-experiment 400 sentinel and BraintrustInvokeError handling for invocation 500s; the new transport raises HTTP errors before callers can inspect raw failure responses.
  • Preserve streaming behavior without restarting a partially consumed response.
  • Preserve custom set_http_adapter() support and caller-owned adapter/session lifecycle. Keep one retry owner per request to avoid multiplying adapter, transport, and writer retries.
  • Preserve credential isolation and organization routing across login changes and loader caches.

Separate /logs3 reliability milestone

Track these behavioral requirements separately from mechanical caller migration, and implement them together with the ingestion migration in PR 1:

  • Honor Retry-After and introduce a thread-safe cooldown shared by batches for the same ingestion destination. New batches must respect an active cooldown too.
  • Keep bounded retry budgets and avoid multiplying transport retries underneath writer retries.
  • Retain/requeue failed batches with bounded backpressure and preserved update ordering; handle permanent failures explicitly instead of retrying indefinitely.
  • Make queue-capacity exhaustion visible. The current LogQueue uses a bounded deque that can evict older records silently, so blindly putting failed records back onto it is insufficient for reliable retention.
  • Make explicit flush() report unsuccessful delivery and preserve useful error details, without requiring BRAINTRUST_SYNC_FLUSH.
  • Cover throttling across concurrent batches, cooldowns longer than the retry budget, retry exhaustion, and existing terminal 413 behavior.

Service-side retention/requeueing of scores when persistence fails is a separate service change. Durable recovery across SDK process crashes is also a separate scope decision. OpenTelemetry's exporter owns its transport and needs a separate assessment if it is to be included; provider SDK networking is outside this migration.

Validation and completion criteria

  • Add failing regressions before behavior fixes, using the existing local HTTP test server for transport behavior and existing cassette-backed coverage where appropriate.
  • Validate each migrated caller group narrowly first, then run test_core, the relevant nox sessions, and static checks from py/ using mise and the authoritative nox/CI configuration.
  • Cover special errors, streaming, uploads, custom adapters, timeouts, credential changes, and connection cleanup.
  • Preserve the distinct LOG_INGESTION mode and validate ingestion performance with the repo's SDK benchmarking workflow, alongside concurrent throttling and recovery tests.
  • No SDK-owned legacy HTTPConnection request path remains after migration, and any retained connection API is explicitly documented for compatibility.
  • Do not claim the throttling/data-loss issue is resolved until the separate ingestion reliability requirements are met.

Source references

Inspected at checkout commit 6537ec61d7809e9c06e8ba9d1955ca6b3eb53cac:

主要语言
Python
星标
21
派生
23
平均合并
20 小时 8 分钟
30 天内合并 PR
94

环境准备

  • 提供 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

braintrustdata/braintrust-sdk-python 的其他 Issue

查看 braintrustdata/braintrust-sdk-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。