Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

A reaper retry decision can overwrite a task that someone else already handled

Aperta
#71 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
64/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
lua, python, redis

Direzione di ricerca

Read reaper.lua, threadmill/backends/lua/acknowledge.lua, and the Redis backend in threadmill/backends/redis.py, then inspect commit 382880b and the TestRedisBrokerReap guard tests. Trace how the claim reaches acknowledge and requeue decisions. Done means stale decisions are dropped while matching-claim acknowledge and requeue paths pass, including the worker-ack, claim-takeover, and inspector-action race tests.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug real

After the lease-expiry retry change (commit 382880b), the reaper hands expired tasks to the retry callback:

  1. reaper.lua claims expired running entries. It renews their lease to RedisBroker.CLAIM_TTL and returns the IDs. It keeps the task data hash.
  2. RedisBroker._reap_task deserializes each claimed task, appends the AcknowledgementTimeout error, and evaluates the retry callback of the task. Then it calls backend.requeue(...) to retry or backend.acknowledge(...) to finalize.

The claim protects against two brokers that claim the same task in the same pass. The select-and-renew step in the script is atomic. But the decision does not make sure that the broker still holds the claim. A stale decision can act on a task that someone else already handled:

  1. Late worker acknowledgement (single node). A slow task finishes after its lease expires. The acknowledge() call of the worker stores SUCCESSFUL and removes the task hash. If the broker read the data before that, its requeue() removes the result and overwrites the task data. It also adds the task to the deferred set again. The task runs again although it succeeded.
  2. Stalled broker, claim taken over. Broker A claims a task and then stalls past CLAIM_TTL (a GC pause, a slow retry callback, or a network problem). Broker B claims the task again and completes the decision. When A starts again, its stale decision overwrites the outcome from B and can schedule the task twice.
  3. Inspector action. A user requeues or removes the task between the claim and the decision. The decision of the broker undoes that action.

The finalize path is mostly protected by the ZREM guard in acknowledge.lua. A second acknowledgement is a no operation. The dangerous operation is mainly requeue, which returns the task to the queue. All paths can also overwrite the task data.

Proposed correction

Make the reap decision conditional on the claim that produced it:

  • Let reaper.lua write a claim identity with the running entry. Use the claim deadline and compare ZSCORE, or use a token in the task hash.
  • Give acknowledge() and requeue() an optional guard parameter. The Lua scripts must make sure that the parameter matches before they write. A mismatch discards the decision. The claim then lapses and the next pass decides again. The result is a delay, not a lost task.
  • Tests: the worker-ack race, claim takeover after CLAIM_TTL, inspector dequeue between claim and decision, and the matching-claim path for both acknowledge and requeue.

An implementation of this guard was written and then removed to keep the lease-expiry retry change small. The code can return from commit 382880b (files threadmill/backends/lua/acknowledge.lua, threadmill/backends/redis.py, and the TestRedisBrokerReap guard tests).

Lingua principale
Python
Stelle
19
Fork
1
Merge medio
9h 46m
PR unite (30g)
17

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di codingjoe/threadmill

Tutte le issue di codingjoe/threadmill

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.