Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

A reaper retry decision can overwrite a task that someone else already handled

Geschlossen
#71 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
64/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
lua, python, redis

Rechercherichtung

Read reaper.lua, threadmill/backends/lua/acknowledge.lua, and the Redis backend in threadmill/backends/redis.py, then inspect commit 382880b and the TestRedisBrokerReap guard tests. Trace how the claim reaches acknowledge and requeue decisions. Done means stale decisions are dropped while matching-claim acknowledge and requeue paths pass, including the worker-ack, claim-takeover, and inspector-action race tests.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

bug real

After the lease-expiry retry change (commit 382880b), the reaper hands expired tasks to the retry callback:

  1. reaper.lua claims expired running entries. It renews their lease to RedisBroker.CLAIM_TTL and returns the IDs. It keeps the task data hash.
  2. RedisBroker._reap_task deserializes each claimed task, appends the AcknowledgementTimeout error, and evaluates the retry callback of the task. Then it calls backend.requeue(...) to retry or backend.acknowledge(...) to finalize.

The claim protects against two brokers that claim the same task in the same pass. The select-and-renew step in the script is atomic. But the decision does not make sure that the broker still holds the claim. A stale decision can act on a task that someone else already handled:

  1. Late worker acknowledgement (single node). A slow task finishes after its lease expires. The acknowledge() call of the worker stores SUCCESSFUL and removes the task hash. If the broker read the data before that, its requeue() removes the result and overwrites the task data. It also adds the task to the deferred set again. The task runs again although it succeeded.
  2. Stalled broker, claim taken over. Broker A claims a task and then stalls past CLAIM_TTL (a GC pause, a slow retry callback, or a network problem). Broker B claims the task again and completes the decision. When A starts again, its stale decision overwrites the outcome from B and can schedule the task twice.
  3. Inspector action. A user requeues or removes the task between the claim and the decision. The decision of the broker undoes that action.

The finalize path is mostly protected by the ZREM guard in acknowledge.lua. A second acknowledgement is a no operation. The dangerous operation is mainly requeue, which returns the task to the queue. All paths can also overwrite the task data.

Proposed correction

Make the reap decision conditional on the claim that produced it:

  • Let reaper.lua write a claim identity with the running entry. Use the claim deadline and compare ZSCORE, or use a token in the task hash.
  • Give acknowledge() and requeue() an optional guard parameter. The Lua scripts must make sure that the parameter matches before they write. A mismatch discards the decision. The claim then lapses and the next pass decides again. The result is a delay, not a lost task.
  • Tests: the worker-ack race, claim takeover after CLAIM_TTL, inspector dequeue between claim and decision, and the matching-claim path for both acknowledge and requeue.

An implementation of this guard was written and then removed to keep the lease-expiry retry change small. The code can return from commit 382880b (files threadmill/backends/lua/acknowledge.lua, threadmill/backends/redis.py, and the TestRedisBrokerReap guard tests).

Vorherrschende Sprache
Python
Sterne
19
Forks
1
Ø Merge
14 Std. 39 Min.
Gemergte PRs (30 T.)
21

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus codingjoe/threadmill

Alle Issues in codingjoe/threadmill

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.