Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

A reaper retry decision can overwrite a task that someone else already handled

未关闭
#71 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
64/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
lua, python, redis

调研方向

Read reaper.lua, threadmill/backends/lua/acknowledge.lua, and the Redis backend in threadmill/backends/redis.py, then inspect commit 382880b and the TestRedisBrokerReap guard tests. Trace how the claim reaches acknowledge and requeue decisions. Done means stale decisions are dropped while matching-claim acknowledge and requeue paths pass, including the worker-ack, claim-takeover, and inspector-action race tests.

由索引模型根据 Issue 内容生成。

描述

bug real

After the lease-expiry retry change (commit 382880b), the reaper hands expired tasks to the retry callback:

  1. reaper.lua claims expired running entries. It renews their lease to RedisBroker.CLAIM_TTL and returns the IDs. It keeps the task data hash.
  2. RedisBroker._reap_task deserializes each claimed task, appends the AcknowledgementTimeout error, and evaluates the retry callback of the task. Then it calls backend.requeue(...) to retry or backend.acknowledge(...) to finalize.

The claim protects against two brokers that claim the same task in the same pass. The select-and-renew step in the script is atomic. But the decision does not make sure that the broker still holds the claim. A stale decision can act on a task that someone else already handled:

  1. Late worker acknowledgement (single node). A slow task finishes after its lease expires. The acknowledge() call of the worker stores SUCCESSFUL and removes the task hash. If the broker read the data before that, its requeue() removes the result and overwrites the task data. It also adds the task to the deferred set again. The task runs again although it succeeded.
  2. Stalled broker, claim taken over. Broker A claims a task and then stalls past CLAIM_TTL (a GC pause, a slow retry callback, or a network problem). Broker B claims the task again and completes the decision. When A starts again, its stale decision overwrites the outcome from B and can schedule the task twice.
  3. Inspector action. A user requeues or removes the task between the claim and the decision. The decision of the broker undoes that action.

The finalize path is mostly protected by the ZREM guard in acknowledge.lua. A second acknowledgement is a no operation. The dangerous operation is mainly requeue, which returns the task to the queue. All paths can also overwrite the task data.

Proposed correction

Make the reap decision conditional on the claim that produced it:

  • Let reaper.lua write a claim identity with the running entry. Use the claim deadline and compare ZSCORE, or use a token in the task hash.
  • Give acknowledge() and requeue() an optional guard parameter. The Lua scripts must make sure that the parameter matches before they write. A mismatch discards the decision. The claim then lapses and the next pass decides again. The result is a delay, not a lost task.
  • Tests: the worker-ack race, claim takeover after CLAIM_TTL, inspector dequeue between claim and decision, and the matching-claim path for both acknowledge and requeue.

An implementation of this guard was written and then removed to keep the lease-expiry retry change small. The code can return from commit 382880b (files threadmill/backends/lua/acknowledge.lua, threadmill/backends/redis.py, and the TestRedisBrokerReap guard tests).

主要语言
Python
星标
19
派生
1
平均合并
9 小时 46 分钟
30 天内合并 PR
17

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

codingjoe/threadmill 的其他 Issue

查看 codingjoe/threadmill 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。