Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

A reaper retry decision can overwrite a task that someone else already handled

Fermée
#71 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
4/5
Temps estimé
3-5 jours
Accessibilité débutants
64/100
Type d'issue
Bug
Clarté
Clairement spécifiée
Activité
Active
Stack technique
lua, python, redis

Piste de recherche

Read reaper.lua, threadmill/backends/lua/acknowledge.lua, and the Redis backend in threadmill/backends/redis.py, then inspect commit 382880b and the TestRedisBrokerReap guard tests. Trace how the claim reaches acknowledge and requeue decisions. Done means stale decisions are dropped while matching-claim acknowledge and requeue paths pass, including the worker-ack, claim-takeover, and inspector-action race tests.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

bug real

After the lease-expiry retry change (commit 382880b), the reaper hands expired tasks to the retry callback:

  1. reaper.lua claims expired running entries. It renews their lease to RedisBroker.CLAIM_TTL and returns the IDs. It keeps the task data hash.
  2. RedisBroker._reap_task deserializes each claimed task, appends the AcknowledgementTimeout error, and evaluates the retry callback of the task. Then it calls backend.requeue(...) to retry or backend.acknowledge(...) to finalize.

The claim protects against two brokers that claim the same task in the same pass. The select-and-renew step in the script is atomic. But the decision does not make sure that the broker still holds the claim. A stale decision can act on a task that someone else already handled:

  1. Late worker acknowledgement (single node). A slow task finishes after its lease expires. The acknowledge() call of the worker stores SUCCESSFUL and removes the task hash. If the broker read the data before that, its requeue() removes the result and overwrites the task data. It also adds the task to the deferred set again. The task runs again although it succeeded.
  2. Stalled broker, claim taken over. Broker A claims a task and then stalls past CLAIM_TTL (a GC pause, a slow retry callback, or a network problem). Broker B claims the task again and completes the decision. When A starts again, its stale decision overwrites the outcome from B and can schedule the task twice.
  3. Inspector action. A user requeues or removes the task between the claim and the decision. The decision of the broker undoes that action.

The finalize path is mostly protected by the ZREM guard in acknowledge.lua. A second acknowledgement is a no operation. The dangerous operation is mainly requeue, which returns the task to the queue. All paths can also overwrite the task data.

Proposed correction

Make the reap decision conditional on the claim that produced it:

  • Let reaper.lua write a claim identity with the running entry. Use the claim deadline and compare ZSCORE, or use a token in the task hash.
  • Give acknowledge() and requeue() an optional guard parameter. The Lua scripts must make sure that the parameter matches before they write. A mismatch discards the decision. The claim then lapses and the next pass decides again. The result is a delay, not a lost task.
  • Tests: the worker-ack race, claim takeover after CLAIM_TTL, inspector dequeue between claim and decision, and the matching-claim path for both acknowledge and requeue.

An implementation of this guard was written and then removed to keep the lease-expiry retry change small. The code can return from commit 382880b (files threadmill/backends/lua/acknowledge.lua, threadmill/backends/redis.py, and the TestRedisBrokerReap guard tests).

Langage dominant
Python
Étoiles
19
Forks
1
Merge moyen
14 h 39 min
PR mergées (30 j)
21

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de codingjoe/threadmill

Toutes les issues de codingjoe/threadmill

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.