Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

A reaper retry decision can overwrite a task that someone else already handled

Abierto
#71 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
64/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
lua, python, redis

Línea de trabajo

Read reaper.lua, threadmill/backends/lua/acknowledge.lua, and the Redis backend in threadmill/backends/redis.py, then inspect commit 382880b and the TestRedisBrokerReap guard tests. Trace how the claim reaches acknowledge and requeue decisions. Done means stale decisions are dropped while matching-claim acknowledge and requeue paths pass, including the worker-ack, claim-takeover, and inspector-action race tests.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug real

After the lease-expiry retry change (commit 382880b), the reaper hands expired tasks to the retry callback:

  1. reaper.lua claims expired running entries. It renews their lease to RedisBroker.CLAIM_TTL and returns the IDs. It keeps the task data hash.
  2. RedisBroker._reap_task deserializes each claimed task, appends the AcknowledgementTimeout error, and evaluates the retry callback of the task. Then it calls backend.requeue(...) to retry or backend.acknowledge(...) to finalize.

The claim protects against two brokers that claim the same task in the same pass. The select-and-renew step in the script is atomic. But the decision does not make sure that the broker still holds the claim. A stale decision can act on a task that someone else already handled:

  1. Late worker acknowledgement (single node). A slow task finishes after its lease expires. The acknowledge() call of the worker stores SUCCESSFUL and removes the task hash. If the broker read the data before that, its requeue() removes the result and overwrites the task data. It also adds the task to the deferred set again. The task runs again although it succeeded.
  2. Stalled broker, claim taken over. Broker A claims a task and then stalls past CLAIM_TTL (a GC pause, a slow retry callback, or a network problem). Broker B claims the task again and completes the decision. When A starts again, its stale decision overwrites the outcome from B and can schedule the task twice.
  3. Inspector action. A user requeues or removes the task between the claim and the decision. The decision of the broker undoes that action.

The finalize path is mostly protected by the ZREM guard in acknowledge.lua. A second acknowledgement is a no operation. The dangerous operation is mainly requeue, which returns the task to the queue. All paths can also overwrite the task data.

Proposed correction

Make the reap decision conditional on the claim that produced it:

  • Let reaper.lua write a claim identity with the running entry. Use the claim deadline and compare ZSCORE, or use a token in the task hash.
  • Give acknowledge() and requeue() an optional guard parameter. The Lua scripts must make sure that the parameter matches before they write. A mismatch discards the decision. The claim then lapses and the next pass decides again. The result is a delay, not a lost task.
  • Tests: the worker-ack race, claim takeover after CLAIM_TTL, inspector dequeue between claim and decision, and the matching-claim path for both acknowledge and requeue.

An implementation of this guard was written and then removed to keep the lease-expiry retry change small. The code can return from commit 382880b (files threadmill/backends/lua/acknowledge.lua, threadmill/backends/redis.py, and the TestRedisBrokerReap guard tests).

Lenguaje dominante
Python
Estrellas
19
Forks
1
Merge medio
9 h 46 min
PR fusionados (30 d)
17

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de codingjoe/threadmill

Todos los issues de codingjoe/threadmill

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.