Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

fix(pipeline): preflight kill-stale 在 v2 console/worker 路径把并发兄弟 run 当孤儿杀——同仓并发 run 互杀成链(今日 4 例实证)

未关闭
#148 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python, redis, rust

调研方向

Start with .dev/flows/self_improve_flow_v2.py, especially _kill_stale_agents() and its live-run detection and parent-process checks; then read test/flow_v2_paths.py and its s33 assertions. Reproduce the concurrent-run case described in the issue and verify the preflight does not kill a sibling run while still terminating a genuine orphan. Keep the existing self-run protection and kill-stale.log format intact.

由索引模型根据 Issue 内容生成。

描述

优先级 P1 · 依赖:无

基线 HEAD = c392d3c1

Summary

self-improve-v2(sbx 变体同源)preflight 的 _kill_stale_agents() 靠「进程 cmdline 里的 --run-id」构造 live run 集(#136 交付的归属过滤):

  • .dev/flows/self_improve_flow_v2.py:158 — live = set([rd.name])(初值只有本次 run)
  • .dev/flows/self_improve_flow_v2.py:160 — m = re.search(r"--run-id[\s=]+(\S+)", c):唯一的外部 live 来源
  • .dev/flows/self_improve_flow_v2.py:179 — if not ids or ids[0] in live: continue:run id 不在 live 即判孤儿
  • .dev/flows/self_improve_flow_v2.py:187 — 父进程白名单只认 self_improve_bridge / self-improve.flow.js
  • .dev/flows/self_improve_flow_v2.py:192 — os.kill(pid, signal.SIGTERM)

这两条 live 通道(带 --run-id 的 bridge 宿主进程、self_improve_bridge 父进程)只在 v1 bridge / flowcast 路径存在。v2 console/worker 路径下 flow 直接在 flow_worker 进程内跑:recursive executor 的父进程是 worker(不在白名单),且全机没有任何进程 cmdline 含 --run-id。

反证命令(本机 Mac worker 现场,13:4x):

$ ps -axo pid,command | grep -- "--run-id" | grep -v grep | wc -l
0

⇒ live 集恒等于 {当前 run} ⇒ 所有并发兄弟 run 的 executor 都被判为「旧 run 孤儿」并 SIGTERM。而它们当时都持有活跃租约(不是真孤儿)。

实证:今日 4 例互杀链(kill-stale.log 的 killed 行与 worker 日志 1:1 对应)
kill-stale 时刻 新 run(日志 live_runs) 被杀 executor 所属 run pid
11:56:09 pipeline-88-1008115608 (#88) pipeline-86-1008091709 (#86) 51584
12:20:24 pipeline-132-1008122023 (#132) pipeline-88-1008115608 (#88) 52320
13:03:30 pipeline-86-1008130329 (#86) pipeline-132-1008122023 (#132) 76023
13:09:35 pipeline-88-1008130935 (#88) pipeline-86-1008130329 (#86) 1938

留痕(每 run 目录内):recursive/.flowcast/runs/<run>/kill-stale.log,例如

# kill-stale 2026-10-08T13:09:35
live_runs=pipeline-88-1008130935
killed pid=1938 pgid=1938 run=pipeline-86-1008130329 cmd=recursive --workspace .../pipeline-86-1008130329/worktree ... run #86 ...

下游因果(每例同构):被杀方 impl 节点立刻以 AgentRunError: executor 'recursive' exited 143 失败(~/.plaita-console/worker-mac.log 13:09:35 与 kill-stale 13:09:35 同秒)→ 节点声明「不终态化,等消息重投后重跑失败节点(第 1/5 次)」→ 但该消息早已被另一 worker 按「租约冲突」ack 释放(plaita#50)⇒ 重投永不到来 ⇒ execution 停在 running、无租约 / 不在 PEL / 无 worker 认领 = 僵尸,占槽至 2h 僵尸线才恢复。

影响

同机同仓并发 ≥2 个 run 时,每个新 run 的 preflight 都会杀掉上一个 run 的 agent:被杀的 run 白跑 + 变僵尸占槽(全局 max_in_flight=6 的一格)+ 需要人工 cancel/reopen 才能释放。今日 #86 / #88 / #132 三单的轮流失败与三连僵尸(值守两次人工处置)全部出自此链。recursive 是 Rust 长任务、按仓配额本来就是 2,等于该配额下的并发能力被这条链自噬。

建议

把「live run 集」从执行面取,而不是从 cmdline 取。三选一(按侵入度排序):

  1. 保守法:父进程链上有存活 flow_worker 的 recursive 进程一律视为 live(现在的白名单只认 v1 bridge 两种 cmdline),即只清「父死了且 run 不在本机任何在跑任务里」的真孤儿;
  2. 锁文件法:preflight 把宿主 pid 写进 rd/run.lock(该文件已存在、内容仅 7 字节),判定时 os.kill(pid, 0) 探活即视为 live;
  3. 租约法:preflight 在 worker 进程内可直接读 Redis,run 是否 live = 其 execution 租约键(plaita:execution:lease:*)是否存在。

验收条件(可断言):

  • 并发不误杀:现场起两个同仓 run(两条真进程即可,复用 test/flow_v2_paths.py 的 s33 现场构造),断言新 run 的 preflight killed none(或 killed 清单不含兄弟 run 的 pid),两条都跑到 committed;
  • 真孤儿照杀:宿主已死的 --workspace <旧 run worktree> / --transcript-out <旧 transcript> 进程仍被 SIGTERM,kill-stale.log 格式不变(killed pid=… pgid=… run=… cmd=…);
  • 自保不回归:本 run 重投时不得杀自己上一投递的子树(现有 startswith(rd) 分支保留);s33 现有断言继续绿。

边界

  • 不处理 plaita#50(节点失败重投载体被租约冲突 ack 吃掉)——那是本条的下游放大器,已单独立单;本条修好后僵尸链的放大器仍在(节点重试仍可能丢载体),但触发源(并发互杀)消失。
  • 临时缓解(值守已执行,可回退):pipeline_repo_limits: {jeffkit/recursive: 1}——限 1 槽即无并发兄弟可杀;本条修复落地后回退为 2。
主要语言
Rust
星标
4
派生
0
平均合并
5 小时 32 分钟
30 天内合并 PR
7

环境准备

  • 提供 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 没有贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

jeffkit/recursive 的其他 Issue

查看 jeffkit/recursive 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。