bug: Partial-failure (207) ingestion response retries the whole batch up to 3x; per-item error details only logged at DEBUG
メンテナーはふだん 1 日以内に返信
評価
この issue はまだ評価されていません。
説明
Describe the bug
LangfuseClient._process_response turns an ingestion 207 response that has a non-empty errors list into raise APIErrors([...]) (request.py#L91-L109).
APIErrors is not a subclass of APIError, so the "do not retry 4xx" check in the score ingestion consumer (isinstance(e, APIError) and 400 <= int(e.status) < 500, score_ingestion_consumer.py#L177-L192) never matches, and backoff.on_exception(Exception, max_tries=3) re-posts the whole batch.
Expected: when most events of a score ingestion batch were accepted (successes) and one was rejected with a per-item 4xx, the batch is not re-sent in full, a permanent per-item 4xx is not retried, and the failing event id/status/message is logged.
Actual: the full batch (including events the server already accepted) is POSTed 3 times, and the only thing logged at ERROR level is the generic canned text for the status code ("Bad request. Please check your request ..."), without the item id or the server's message. The server message is only visible at DEBUG level (parse_error.py#L75-L99).
Impact is modest: duplicate delivery is likely absorbed by server-side idempotency on event id (I have not verified this against a server), but it wastes traffic, triples the time the consumer thread spends on the batch (which flush() waits for), and retries errors that can never succeed.
Steps to reproduce
Uses httpx.MockTransport, no server needed (run from a checkout with the repo on PYTHONPATH):
import httpx, json
from queue import Queue
from langfuse._utils.request import LangfuseClient
from langfuse._task_manager.score_ingestion_consumer import ScoreIngestionConsumer
import time; time.sleep = lambda s: None # skip backoff waits
calls = []
body = {"successes": [{"id": "e0", "status": 201}, {"id": "e2", "status": 201}],
"errors": [{"id": "e1", "status": 400, "message": "Invalid request data", "error": "bad"}]}
def h(req):
calls.append(json.loads(req.content)["batch"]); return httpx.Response(207, json=body)
cl = LangfuseClient("pk", "sk", "http://x", "1", 5, httpx.Client(transport=httpx.MockTransport(h)))
q = Queue(); c = ScoreIngestionConsumer(ingestion_queue=q, identifier=0, client=cl, public_key="pk", flush_at=3)
for i in range(3): q.put({"id": f"e{i}", "type": "score-create", "body": {"n": i}})
c.upload()
print("POST count:", len(calls), "ids per post:", [[e["id"] for e in b] for b in calls])
Observed output:
API errors occurred: Bad request. Please check your request for any missing or incorrect parameters. Refer to our API docs: https://api.reference.langfuse.com for details.
POST count: 3 ids per post: [['e0', 'e1', 'e2'], ['e0', 'e1', 'e2'], ['e0', 'e1', 'e2']]
Variant (re-run on the same commit): if the 207 body is malformed ({"errors": ["boom"]} or a JSON list), the .get calls in the 207 branch raise AttributeError (only JSONDecodeError is caught there), and that is also retried 3 times (3 POSTs observed).
Langfuse Cloud or self-hosted?
Not server dependent (reproduced with a mock transport). The 207 shape is the one returned by /api/public/ingestion.
If self-hosted, what version are you running?
n/a
SDK and integration versions
langfuse 4.17.0 (main @ bf11ec121479145923ff27b32ec2ad1d06fd2d51), Python 3.14.0, httpx and backoff as pinned in uv.lock.
Additional information
Possible direction: handle APIErrors separately in _upload_batch: log each item (id, status, message) at warning/error level, treat per-item 4xx other than 429/408 as final, and optionally re-send only the items that failed with 429/5xx. A smaller change is to not retry an APIErrors whose statuses are all non-retryable 4xx. The 207 branch could also check isinstance(payload, dict) and the entries before calling .get.
Related but different: #1874 (log non-retryable 4xx batch drops) covers a plain 4xx response, not the 207 path.
Are you interested in contributing a fix for this bug?
Yes
- 主要言語
- Python
- スター
- 498
- フォーク
- 361
- 平均マージ
- 16時間 40分
- マージ済み PR(30日)
- 38
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
langfuse/langfuse-python のほかの issue
-
bug feat-prompt-management sdk-python
難易度 2/5 1〜3時間 初心者へのやさしさ 80/100
langfuse/langfuse-python#1976 ·
メンテナーはふだん 1 日以内に返信
-
bug: ChatPromptClient.compile appends str(whole list) once per non-dict placeholder item対応中かも @hassiebp が 2 日前に担当しました。 オープンbug feat-prompt-management sdk-python
langfuse/langfuse-python#1971 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
bug: LANGFUSE_MAX_EVENT_SIZE_BYTES is parsed but never enforced for score events対応中かも @hassiebp が 2 日前に担当しました。 オープンbug feat-scores sdk-python
langfuse/langfuse-python#1966 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
bug: non-numeric status in a 207 error entry raises inside handle_exception, kills the score consumer thread, flush()/shutdown() then hang対応中かも @hassiebp が 2 日前に担当しました。 オープン[Integrations] Language Clients bug feat-scores sdk-python
langfuse/langfuse-python#1967 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
bug: GCS upload detection in MediaManager uses substring match on the full URL対応中かも @hassiebp が 10 日前に担当しました。 オープンbug feat-multimodal-media sdk-python
langfuse/langfuse-python#1913 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
langfuse/langfuse-python の issue をすべて見る
似ている issue
-
namespace operations
難易度 1/5 1時間未満 初心者へのやさしさ 72/100
EclipseFdn/open-vsx.org#14043 ·
メンテナーはふだん 1 日以内に返信
-
netbox status: needs triage type: bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
netbox-community/netbox#23376 ·
メンテナーはふだん 1 日以内に返信
-
feedback simulation workshop
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
githubnext/gh-aw-workshop#4455 ·
メンテナーはふだん 1 日以内に返信
-
Triage 🩺
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitオープンneeds-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 77/100
krkn-chaos/krkn#1627 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信