bug: Partial-failure (207) ingestion response retries the whole batch up to 3x; per-item error details only logged at DEBUG
I maintainer di solito rispondono entro 1 giorno
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Describe the bug
LangfuseClient._process_response turns an ingestion 207 response that has a non-empty errors list into raise APIErrors([...]) (request.py#L91-L109).
APIErrors is not a subclass of APIError, so the "do not retry 4xx" check in the score ingestion consumer (isinstance(e, APIError) and 400 <= int(e.status) < 500, score_ingestion_consumer.py#L177-L192) never matches, and backoff.on_exception(Exception, max_tries=3) re-posts the whole batch.
Expected: when most events of a score ingestion batch were accepted (successes) and one was rejected with a per-item 4xx, the batch is not re-sent in full, a permanent per-item 4xx is not retried, and the failing event id/status/message is logged.
Actual: the full batch (including events the server already accepted) is POSTed 3 times, and the only thing logged at ERROR level is the generic canned text for the status code ("Bad request. Please check your request ..."), without the item id or the server's message. The server message is only visible at DEBUG level (parse_error.py#L75-L99).
Impact is modest: duplicate delivery is likely absorbed by server-side idempotency on event id (I have not verified this against a server), but it wastes traffic, triples the time the consumer thread spends on the batch (which flush() waits for), and retries errors that can never succeed.
Steps to reproduce
Uses httpx.MockTransport, no server needed (run from a checkout with the repo on PYTHONPATH):
import httpx, json
from queue import Queue
from langfuse._utils.request import LangfuseClient
from langfuse._task_manager.score_ingestion_consumer import ScoreIngestionConsumer
import time; time.sleep = lambda s: None # skip backoff waits
calls = []
body = {"successes": [{"id": "e0", "status": 201}, {"id": "e2", "status": 201}],
"errors": [{"id": "e1", "status": 400, "message": "Invalid request data", "error": "bad"}]}
def h(req):
calls.append(json.loads(req.content)["batch"]); return httpx.Response(207, json=body)
cl = LangfuseClient("pk", "sk", "http://x", "1", 5, httpx.Client(transport=httpx.MockTransport(h)))
q = Queue(); c = ScoreIngestionConsumer(ingestion_queue=q, identifier=0, client=cl, public_key="pk", flush_at=3)
for i in range(3): q.put({"id": f"e{i}", "type": "score-create", "body": {"n": i}})
c.upload()
print("POST count:", len(calls), "ids per post:", [[e["id"] for e in b] for b in calls])
Observed output:
API errors occurred: Bad request. Please check your request for any missing or incorrect parameters. Refer to our API docs: https://api.reference.langfuse.com for details.
POST count: 3 ids per post: [['e0', 'e1', 'e2'], ['e0', 'e1', 'e2'], ['e0', 'e1', 'e2']]
Variant (re-run on the same commit): if the 207 body is malformed ({"errors": ["boom"]} or a JSON list), the .get calls in the 207 branch raise AttributeError (only JSONDecodeError is caught there), and that is also retried 3 times (3 POSTs observed).
Langfuse Cloud or self-hosted?
Not server dependent (reproduced with a mock transport). The 207 shape is the one returned by /api/public/ingestion.
If self-hosted, what version are you running?
n/a
SDK and integration versions
langfuse 4.17.0 (main @ bf11ec121479145923ff27b32ec2ad1d06fd2d51), Python 3.14.0, httpx and backoff as pinned in uv.lock.
Additional information
Possible direction: handle APIErrors separately in _upload_batch: log each item (id, status, message) at warning/error level, treat per-item 4xx other than 429/408 as final, and optionally re-send only the items that failed with 429/5xx. A smaller change is to not retry an APIErrors whose statuses are all non-retryable 4xx. The 207 branch could also check isinstance(payload, dict) and the entries before calling .get.
Related but different: #1874 (log non-retryable 4xx batch drops) covers a plain 4xx response, not the 207 path.
Are you interested in contributing a fix for this bug?
Yes
- Lingua principale
- Python
- Stelle
- 498
- Fork
- 361
- Merge medio
- 16h 40m
- PR unite (30g)
- 38
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di langfuse/langfuse-python
-
bug feat-prompt-management sdk-python
Difficoltà 2/5 1-3 ore Idoneità per principianti 80/100
langfuse/langfuse-python#1976 ·
I maintainer di solito rispondono entro 1 giorno
-
bug: ChatPromptClient.compile appends str(whole list) once per non-dict placeholder itemForse già presa @hassiebp l’ha presa 2 giorni fa. Apertabug feat-prompt-management sdk-python
langfuse/langfuse-python#1971 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug: LANGFUSE_MAX_EVENT_SIZE_BYTES is parsed but never enforced for score eventsForse già presa @hassiebp l’ha presa 2 giorni fa. Apertabug feat-scores sdk-python
langfuse/langfuse-python#1966 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug: non-numeric status in a 207 error entry raises inside handle_exception, kills the score consumer thread, flush()/shutdown() then hangForse già presa @hassiebp l’ha presa 2 giorni fa. Aperta[Integrations] Language Clients bug feat-scores sdk-python
langfuse/langfuse-python#1967 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug: GCS upload detection in MediaManager uses substring match on the full URLForse già presa @hassiebp l’ha presa 10 giorni fa. Apertabug feat-multimodal-media sdk-python
langfuse/langfuse-python#1913 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di langfuse/langfuse-python
Issue simili
-
Claiming namespace `jft63`Apertanamespace operations
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
EclipseFdn/open-vsx.org#14043 ·
I maintainer di solito rispondono entro 1 giorno
-
netbox status: needs triage type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
netbox-community/netbox#23376 ·
I maintainer di solito rispondono entro 1 giorno
-
feedback simulation workshop
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
githubnext/gh-aw-workshop#4455 ·
I maintainer di solito rispondono entro 1 giorno
-
Triage 🩺
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitApertaneeds-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 77/100
krkn-chaos/krkn#1627 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno