tracking: LLM 适配器线协议与 usage(3 项)—— 坏请求/坏数值被固化
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 15/100
Research direction
This is a tracking issue that produces no code itself; the work sits in sub-issues #163, #157 and #168, which must land in order. Start with the usage-normalization sub-issue #163: read the usage mapping in src/llm/anthropic.rs, the threshold check in src/compact/mod.rs, and the cost logic in src/llm/pricing.rs. Done means the acceptance items listed in the issue pass, such as prompt_tokens matching total input in the adapter test near lines 1530-1540.
Written by the indexing model from the issue text.
Description
优先级 P0 · 依赖:无 · 类型:tracking(簇设计入口,不产出代码;子单全部落地后关闭)
记号:正文中的「不变量 #N」指.dev/AGENTS.md编号,不是 issue 编号;#150这类才是 issue。本簇子单头部写depends-on: #N表示落地串行(同文件/同契约),不会与语义阻塞混淆。
基线 HEAD = 14bac15a(子单同基线)。本单来自一次全仓审计:先按根因归簇,再拆成可独立验收的子单。
共同根因
三条都在适配器边界上,且都会把一次坏请求/坏响应固化成长期故障:
- usage 未归一化(#163):Anthropic 的
input_tokens只算未命中缓存的部分,而适配器只把它赋给prompt_tokens(src/llm/anthropic.rs:663-667、:1314-1324;适配器自己的测试:1530-1540就显示「总输入 150、prompt_tokens=100」)。后果两条:src/compact/mod.rs:139-143只要actual > 0就用 token 阈值并完全跳过字节回退 → 缓存开启后主动压缩永不触发,直到撞上下文 400;src/llm/pricing.rs:198-210在cache_hit_tokens == 0时按prompt_tokens计费 → 首次请求(缓存写 = 整段前缀)完全不计成本,命中时缓存写也只按输入价(实际 1.25×)。 - 请求里出现非首位 system 角色(#157):
extract_system_message只看messages.first()(anthropic.rs:948-956),其余 System 由serialize_message输出"role":"system"(:1146-1151)。Anthropic 只接受user/assistant。三种触发:deferred tools 块被 prepend 在系统提示之前(src/run_core.rs:1180-1182,api.anthropic.com默认满足supports_deferred_tools,而TodoWrite恒注册)→ 每次请求都 400;压缩后重注入的 System 附件(src/runtime.rs:895-969);技能 glob 注入(run_core.rs:795-800)。现有tests/deferred_tool_loading.rs用不校验 role 的 mock,测不出。 - 中断落盘空 assistant 消息(#168):
have_partial把「只有 tool_call 增量、无文本」也算数(src/llm/openai.rs:851-866、anthropic.rs:555-569),随后把 tool_calls 清空(openai.rs:886-902)并 push 一条空 content 的 assistant 消息(src/run_core.rs:1716-1717);Anthropic 序列化会发出空文本块(anthropic.rs:1183-1187),而 #16 的空 content 守卫只在tool_calls非空时生效(openai.rs:1240)→ 此后每一轮都 400。
子单与依赖
| 子单 | 依赖 | 范围 |
|---|---|---|
| #163 | 无 | usage 归一化(prompt = input + cache_read + cache_creation)+ pricing 按 hit/miss 计费(含 cache-write 倍率) |
| #157 | depends-on: #163 | 所有 System 角色进顶层 system(或转 user reminder);deferred 块放到系统提示之后 |
| #168 | depends-on: #157 | 只在有文本时保留半成品;空 content 且无 tool_calls 的 assistant 消息在落盘与序列化两侧都跳过 |
落地思路
- 顺序有讲究:#163 不动请求形状,先落它可以让 A 簇的压缩触发判据(#162 → #158 → #156)建立在正确数值上;#157 改请求形状,需同步改
tests/deferred_tool_loading.rs的断言与新增「无messages[*].role == system」的形状测试;#168 跨openai.rs/anthropic.rs/run_core.rs,放最后避免与前两者抢同一处序列化代码。 - #163 一处改动两处受益(压缩触发 + 计费),若评审要求拆单,可拆 usage 归一化与 pricing 费率两单,但必须同批落地——只改一处会造成重复计费或漏计。
pricing.rs的改动要与归一化后的prompt_tokens语义对齐(保持「hit + miss == 总输入」),并给ModelPricing预留 cache-write 倍率字段。- Anthropic 的 400 结论基于官方 role 约束 + 序列化代码路径,未打真 API 验证(无凭据);#157 建议在验收里要求:新增请求形状测试 + 有条件时打一次真实端点。
簇级验收
anthropic.rs:1530-1540的期望改为prompt_tokens == 总输入;cache_read=0, cache_creation=40_000, input=50时成本不为 0;缓存场景下should_compact能按真实上下文触发;- 序列化后
messages[*].role只出现user/assistant,System 文本全在顶层system;tests/deferred_tool_loading.rs增加 role 断言; - tool_call 阶段断流后 transcript 无空 assistant 消息,后续构造的请求体不含空文本块。
边界
不含 #90 已落地的 prompt cache 开启策略本身;不含其它 provider(OpenAI 兼容侧)的 usage 语义——若发现同型问题,另开单。
实施顺序、PR 拆分与验收矩阵(补充 · 2026-10-10)
顺序(必须串行,前两条都改 src/llm/anthropic.rs)
| 波 | 子单 | 理由 |
|---|---|---|
| W1 | #163 | 只改 usage 数值与计费,不动请求形状;先落它,A 簇 #162 → #158 → #156 的压缩触发才建立在正确数值上 |
| W2 | #157 | 改请求形状(role 与 system 位置),需要同步改 tests/deferred_tool_loading.rs 的断言 |
| W3 | #168 | 跨 openai.rs/anthropic.rs/run_core.rs,且要处理"空 content 消息"这一新判据,放最后避免与 #157 抢同一处序列化代码 |
与 A 簇的交互(重要):#163 与 A 簇 #162 是"输入正确性"的两半——#162 保证用单次读数,#163 保证读数本身包含缓存 token。两单落地前不要调 default_compact_threshold_tokens。
跨簇文件重叠:src/llm/anthropic.rs 被 #163/#157/#168 共用(同簇内串行即可);src/runtime.rs 不在本簇范围(#163 只改适配器与 pricing)。
PR 拆分:#163 若被要求拆成"usage 归一化"与"pricing 费率",必须同一批合并(只改一处会造成重复计费或漏计),建议单 PR 两 commit。
簇级验收 → 测试映射
| 验收 | 落点 |
|---|---|
prompt_tokens == 总输入 |
改 src/llm/anthropic.rs:1530-1540 的期望 |
| 冷缓存步成本 ≠ 0 | src/llm/pricing.rs 新增单测 |
| 缓存场景能触发压缩 | src/compact/mod.rs 或 runtime 级用例(结合 A 簇 #162 的 runtime 测试) |
无 messages[*].role == "system" |
新增请求形状测试(anthropic.rs 单测 + tests/deferred_tool_loading.rs 加 role 断言) |
| 中断后无空 assistant 消息 | src/run_core.rs/tests/ 新增(mock 流先发 tool_call 增量再断流) |
- Dominant language
- Rust
- Stars
- 4
- Forks
- 0
- Avg merge
- 5h 32m
- Merged PRs (30d)
- 7
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from jeffkit/recursive
-
needs-human
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
jeffkit/recursive#159 · 5 comments ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 53/100
jeffkit/recursive#195 · 4 comments ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 15/100
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 8/100
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 12/100
Maintainers usually reply within 1 day
All issues in jeffkit/recursive
Similar issues
-
llm translation
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
area: backend type: enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 61/100
armadavalor/WinKnife#6 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
5omeOtherGuy/phaseone#870 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 72/100
MattA-Official/vwmcp#20 ·