Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Kafka flush timeout returns 202 with delivery unconfirmed

Aperta
#220 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 2 giorni

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
3/5
Tempo stimato
1-2 giorni
Idoneità per principianti
70/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Inizia in src/writers/writer_kafka.py, in KafkaWriter.write() e nel ramo terminale di timeout del flush che emette un avviso e poi raggiunge il flusso di successo; quindi segui come viene registrato il successo in _write_to_all() per POST /topics/{topic_name}. Esamina adr/002-observability/002-observability.md per il contratto richiesto sulla semantica di timeout rispetto a quella degli errori. Esegui o estendi gli unit test del writer Kafka per il caso di timeout remaining > 0 senza eccezione e verifica che il caso non venga più trattato come riuscito (Kafka non viene segnalato in writers_ok).

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug
Describe the bug

KafkaWriter.write() (src/writers/writer_kafka.py) retries the flush, and when messages are still pending after the last attempt it logs a WARNING and falls through:

if isinstance(remaining, int) and remaining > 0:
    logger.warning("Kafka flush timed out with messages still pending.", extra={...})

A pure timeout appends nothing to errors and raises nothing, so the if errors: branch is never reached and write() returns normally. _write_to_all() counts Kafka in writers_ok and the request returns 202, while the message may never have been delivered.

Only a WARNING records it, so an alarm on level = "ERROR" is blind to this by construction — and ADR-002's count(ERROR) == count(5xx) invariant holds here only because the request is never considered failed at all.

Steps to Reproduce
  1. Configure a topic with a Kafka writer and a reachable broker, so the producer initializes and produce() succeeds.
  2. Make the broker unable to acknowledge the write while keeping the connection alive — e.g. take the partition leader offline after produce, or set min.insync.replicas above the number of live in-sync replicas.
  3. POST /topics/{topic_name} with a valid message.
  4. flush() returns remaining > 0 on all 3 attempts (KAFKA_FLUSH_RETRIES, 7s timeout each per KAFKA_FLUSH_TIMEOUT) and raises no KafkaException.
  5. Observe the response: 202, with kafka listed in writers_ok.
  6. Observe the logs: one WARNING ("Kafka flush timed out with messages still pending."), zero ERROR.
Expected state

A message whose delivery was never confirmed must not be reported to the caller as written.

202 is meant to mean "accepted by every configured sink". Either the response reflects that the Kafka write did not complete, or writers_ok stops listing a sink whose delivery is unconfirmed — but a caller must not be told the message landed when the service does not know that it did.

Impact / Severity

High

Attachments / Evidence

src/writers/writer_kafka.py — flush loop, terminal timeout branch, and the if errors: gate that the timeout path never reaches:

        # Warn if messages still pending after retries
        if isinstance(remaining, int) and remaining > 0:
            logger.warning(
                "Kafka flush timed out with messages still pending.",
                extra={"pending_messages": remaining, "flush_timeout_sec": _KAFKA_FLUSH_TIMEOUT_SEC},
            )

        duration_ms = round((time.perf_counter() - started_at) * 1000, 2)

        if errors:                      # <- empty on a pure timeout
            ...
            raise WriteError(failure_text)

        logger.debug("Kafka accepted the message.", ...)   # <- reached instead

errors is appended to only by delivery_report (on a delivery error) and by the two except KafkaException blocks. A flush that simply does not drain in time hits none of them.

Related / References

Two options, and the choice is a product decision rather than a cleanup:

  1. Treat a terminal flush timeout as a WriteError. Correct on the contract, but it turns a current 202 into a 500, so it is a caller-visible behaviour change.
  2. Keep the 202 and alarm on this specific WARNING, treating unconfirmed delivery as an operational signal rather than a request failure.

Option 1 is the honest one if 202 is meant to mean "accepted by every configured sink". Whichever is chosen, the decision belongs in ADR-002 and writers_ok must stop reporting an unconfirmed sink as OK.

Acceptance:

  • Decision recorded in ADR-002 §Logging strategy.
  • writers_ok no longer reports a sink whose delivery was never confirmed.
  • Unit test covers flush timeout with remaining > 0 and no raised exception.

Related: ADR-002 (adr/002-observability/002-observability.md), #193, PR #204, #219 — the other invariant gap found in the same review.

Lingua principale
Python
Stelle
4
Fork
0
Merge medio
1g 20h
PR unite (30g)
8

Preparare l'ambiente

  • Include un Dockerfile o un file Docker Compose
  • Ha un modello di pull request
  • Nessuna guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di AbsaOSS/EventGate

Tutte le issue di AbsaOSS/EventGate

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.