compaction: route summaries through the roles ladder, and cut their cost
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
Hướng nghiên cứu
Start with internal/session/compact_summary.go and internal/session/auxiliary.go, then read the roles ladder under internal/roles and the long-context sandbox procedure from #1658. Run codeaf chat --debug to measure summary and post-pass costs before choosing among the open questions. Done means the accepted items preserve the current default, expose the selected model and cost behavior, pass the stated summary checks, and update internal/manual/chat/compacting-over-and-over.md.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Follow-up to #1658, which added summaries as compaction's last step. Stubbing and folding cost nothing; nearly all of a compaction's cost is two things:
- The summary request, which reads the whole region it replaces.
- The first request after a pass, which misses the prompt cache from the point the pass rewrote.
The proposal: route the summary through codeaf's existing roles ladder, so the model that writes it is a setting like every other errand's. Then, per summary, pick the cheaper of two ways to write it, using the prices the model catalog already carries:
- on a cheaper role model;
- on the conversation's own model, reusing its prompt cache.
Terms
P_in,P_cache,P_out: the conversation model's price per input token, cached input token and output token.P_cacheis often aboutP_in / 10.P_in′,P_out′: the same for the model thesummaryrole resolves to.R: the tokens a summary reads (the region it replaces).A: the tokens the summary answer is (at mostsummaryAnswerTokens, 4,096 on large windows).C: the conversation's size right after a pass.
The model catalog already carries prompt_price, cache_read_price and completion_price per model, so every figure below can be computed before a request is sent.
1. A summary role on the roles ladder (central)
Now: askForSummary (internal/session/compact_summary.go) calls completeWithModel with the conversation's model directly. There's no way to choose another model.
Proposed: register a summary role (internal/roles, e.g. roles.Register(RoleSummary, TierLow)) and send the summary through callRole (internal/session/auxiliary.go), the door titles, captions and task names already use. roles.Ladder then resolves, in order:
- the pin,
roles.summaryin settings; - the tier,
tiers.low(the small work row); - the conversation's model, the ladder's floor.
What callRole brings with no new machinery:
- One fall-through rung: a cheap model that fails or times out hands over to the next rung, usually the conversation's model, instead of leaving the pass without a summary.
- A time limit per tier (
roles.PatienceFor), billing of the model that actually answered, and--one-modelrespected (its floor is the conversation's model). - Route health (
CrewHealthySend) and the crew day cap, which already apply to every helper call throughcompleteWithNamedModel.
Saving when a cheaper model writes it: R × (P_in − P_in′) + A × (P_out − P_out′) per summary, one to two orders of magnitude on expensive conversation models.
2. A cache-reusing request, for when the ladder lands on the conversation's model
Now: the summary request is new text: its own system prompt, then the region re-typed as PERSON: / assistant lines (summaryLines). Nothing matches the cache, so the region costs R × P_in.
Proposed: when the summary is written by the conversation's own model, send the conversation exactly as the last request did (same system prompt, tool definitions and messages), then add one final message asking for a summary of everything before message N.
Saving: about R × (P_in − P_cache), roughly 90% of the summary's input when P_cache ≈ P_in / 10.
Care needed:
- Tool definitions must be sent for the prefix to match, but must not be called:
tool_choice: none, plus a refusal path if a call comes back anyway. - Only the first chunk of a region bigger than one request matches the cache.
- The rolling "Summary so far" must still work.
3. Choosing between 1 and 2, per summary
The two savings are exclusive: the cache only helps the model that already holds it. For each summary, compare:
- Cache reuse on the conversation's model: about
R × P_cache + A × P_out. - The role's cheaper model: about
R × P_in′ + A × P_out′.
Pick the lower one when the role resolves to a different model. When the conversation model's cached rate is below the cheap model's full rate, cache reuse wins even though the model is more expensive.
4. The cache break after a fold
Now: a fold rewrites the middle of the conversation, so the next request costs about C × P_in instead of C × P_cache: an extra C × (P_in − P_cache) per pass. The automatic pass folds only down to its target (about 77% of the window on large models), so C stays large and this can cost as much as a summary.
Proposed: once a pass is going to break the cache anyway, weigh going deeper. A summary makes C small, and every later request then saves about (C_fold − C_summary) × P_cache. This changes the automatic policy, so it waits for measurements (below).
5. A smaller summary answer
summaryAnswerTokens caps A at min(4096, max(512, window/20)). Lowering it saves at most (4096 − new cap) × P_out′ per summary, small next to 1–4.
Catches to design around
- The default must not move. #1658 deliberately has the conversation's own model write summaries, but
tiers.lowalways has a model set, so registering the role as-is would switch every install to the small-work model. The ladder needs an opt-in (see open questions). - The window. Chunking is sized to the conversation model's window (
plan.window). A cheaper model with a smaller window needs chunks sized to its window (ContextWindowFor), and its answer cap re-derived. - Fall-through doubles the bill on a bad minute. A cheap rung that fails after reading the region, then the conversation-model rung, means paying for
Rtwice. This is the same trade every errand makes withroleFallThroughs = 1. - Cache hits aren't guaranteed. The router may send the summary request to a different endpoint (lane) than the one holding the cache, and some providers don't cache at all. The choice in 3 is an estimate until the call's cached-token figure comes back.
Open questions
- Opt-in shape: pin only (
roles.summaryset means use the ladder), or a settings row (e.g. "summaries: conversation model / small work / pick"), or use the ladder by default with the tier rung skipped unless opted in? - Which tier:
TierLow(small work, already used for titles and digests),TierReflex(cheaper still, but read every turn and small-windowed), or a new tier just for summaries? - Automatic or explicit choice: should codeaf pick between cache reuse and the cheaper model by price (item 3), or always use what the person set? If automatic, where does it say which it chose: the
⚭line,/status, the usage record? - Quality floor: today a summary must be prose, not repeat itself, and be smaller than what it replaces. Is that enough for a much smaller model, or does a cheap summary need more checks (e.g. the three kept messages' topics mentioned)?
- Same policy for
/compactand the automatic pass? A person's/compactmight justify the conversation's own model even when a cheaper one is set, or the reverse.
Two narrower ones:
- Refusal recovery: should it use the ladder too? It needs a summary fast and reliably, which argues for the conversation's own model.
- Item 4: what measured break-even (summary cost vs requests saved afterwards) would justify changing the automatic target?
Measure first
- Use the long-context sandbox from #1658's testing (conversations seeded just past the automatic line) with
codeaf chat --debug. - For one automatic summary pass and one fold-only pass, record for the summary request and the first request after the pass: input, cached input and output tokens, and their cost.
- Repeat after items 1–2 land, once with the role pinned to a cheap model and once unpinned.
Acceptance
- Item 1: with nothing set, summary requests are byte-for-byte what they are today (the conversation's model). With the role pinned, the summary's usage line names the pinned model. If that model fails, the summary lands from the next rung.
- Item 2: on a caching provider, the
--debugcall record of a summary request shows cached input covering at least 90% of the region. The summary still splices exactly as today: region unchanged, smaller than what it replaces, prose. - Item 3: the choice between 1 and 2 is made from catalog prices and is visible somewhere a person can check. The open questions say where.
- Items 4 and 5: decided from the measurements, with the figures recorded here.
- Manual:
internal/manual/chat/compacting-over-and-over.mdsays which model writes a summary and how to change it (the manual law).
- Ngôn ngữ chính
- Go
- Star
- 115
- Fork
- 14
- Merge trung bình
- 9 giờ 3 phút
- Pull request đã merge (30 ngày)
- 517
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của Agent-Field/CodeAF
-
area:session bug sev:papercut
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Agent-Field/CodeAF#1779 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
tasks: a stopped task's commit is signed with the chat model, not the model that did the workĐang mởarea:session bug sev:papercut
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Agent-Field/CodeAF#1768 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:chat bug sev:papercut
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Agent-Field/CodeAF#1742 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:build bug sev:papercut
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Agent-Field/CodeAF#1679 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:session bug sev:serious
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Agent-Field/CodeAF#1678 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của Agent-Field/CodeAF
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
vanderheijden86/b9s#20 ·
-
go-battery needs an ndsctl on PATH: TestPurchaseSessionGuardHoldsThroughTheOutcomeUnknownWindow fails on bare hosts (passes with stub)Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
OpenTollGate/tollgate-module-basic-go#726 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
ux waiting for feedback
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 63/100
evcc-io/evcc#34527 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
phase:v3 type:harness
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày