Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

fix(worker): 低盘宿主抢单白跑——claim 无本机盘预检,preflight 拒后退避升档至 6h(实测一夜 11 次)

Open
#49 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python, redis

Research direction

Start with plaita/server/task_queue.py (read() and xreadgroup) and plaita/server/flow_worker.py (_consume_loop); then inspect the flow preflight in recursive/.dev/flows/self_improve_flow_v2.py and its _sbx variant. Check how RECURSIVE_MIN_FREE_DISK_GIB is configured and add focused tests mocking statvfs to verify low-disk workers do not claim while healthy workers do. Done means the listed acceptance scenarios pass, including automatic recovery after disk space returns.

Written by the indexing model from the issue text.

Description

优先级 P2 · 依赖:无

现象:多 worker 池下,磁盘低于守线的 worker 会「抢单白跑」——先领取任务,flow 的 preflight 才判本机盘不足 → retry-later;盘充足的健康 worker 全程空闲却拿不到任务。2026-10-08 00:48–01:05(CST)实测 11+ 次连续白跑。

证据(远端 ~/.issue-keeper/pipeline/runs.jsonl,全部 stage=preflight、disk 15.6–16.3GiB < min;exec 号可与 VM worker 日志 ~/.plaita-console/worker-tcg1.log 的消费记录对上):

  • 00:48:30 recursive#86 75c37fb2 / 00:48:32 agentproc#17 0b1eb8a2 / 00:48:33 agentproc#13 bd7a0918 / 00:48:35 agentproc#8 90e60d87
  • 00:55:00 recursive#87 98c394fe / 00:55:05 recursive#132 bb73a17f / 00:55:10 argusai#11 07d2317b
  • 01:01:24 recursive#88 4ad7247b / 01:01:26 agentproc#17 545984ad / 01:01:28 plaita#24 57f030f8 / 01:01:31 plaita#23 8b96e4ed
  • 对照:同时刻 Mac worker 完全空闲(队列 XLEN=0、PEL=0,本机盘 198GiB 富余)。
  • 退避放大:keeper state 中 #86/#87/#88 的 retry_after 已升至 ≈350min(6h 档;retry_later_streak 6–7)——瞬时低盘被放大成数小时停机。

根因(2026-10-08 快照):

  1. worker 领取任务无本机资源预检:plaita/server/task_queue.py:171 read()(:199 xreadgroup)由 plaita/server/flow_worker.py:2390 _consume_loop 驱动,谁先读谁拿走;
  2. 盘门在 claim 之后才拦:flow 的 preflight 节点 recursive/.dev/flows/self_improve_flow_v2.py:209(sbx 变体 self_improve_flow_v2_sbx.py:198 同)在 run() 里 os.statvfs(repo) → free_gib < RECURSIVE_MIN_FREE_DISK_GIB(默认 20)→ 返回 {"ok": False, "why": "disk … < min"}——此节点跑在领取任务的宿主上。

影响:①每次白跑一轮(claim→preflight→retry-later 回评→重排),issue 交付整体后移 ≈退避时长;②退避升档把可自愈的瞬时低盘放大到 6h;③池内算力错配:健康机闲置、低盘机空转;④每单一条 retry-later 回评(噪声)。

建议修法(最小 → 完整):

  1. worker claim 前预检本机盘:低于 RECURSIVE_MIN_FREE_DISK_GIB 时不 XREADGROUP(空转/短退避轮询),盘回线自动恢复——单点改动即可消除白跑;
  2. (可选)dispatcher 侧按 worker 上报资源路由(需要 worker 心跳携带盘/负载);
  3. (配套)retry-later 退避档位对「宿主资源类」原因(disk/负载)不升档,或在盘回线时对这类挂账重置退避(现 #86/#87/#88 需等 ~6h 或人工 reopen)。

验收条件:

  • 构造「A 机盘 < min、B 机充足」:A worker 不再领取任务(不产生 claim);runs.jsonl 不再出现因 A 宿主低盘导致的 retry-later;同一队列 B 正常领取并完成;A 盘恢复后无需重启自动恢复领取。
  • 单测:mock statvfs/阈值,验证 claim 门(A 不读、B 读)。
Dominant language
Python
Stars
0
Forks
1
Avg merge
3h 53m
Merged PRs (30d)
3

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from jeffkit/plaita

All issues in jeffkit/plaita

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.