Cascade routing mode: user-selected primary model with verifier-gated frontier escalation
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- typescript
- 领域
- ai-infra-agents, cli
调研方向
从子任务 #1138、#1135、#1136 和 #1137 开始,然后跟踪现有的 config、flags 和交互式模型选择入口。检查 verifier stack 如何报告 build、data-test 和 schema-contract 的结果。当用户可以选择两个模型、失败的验证会携带上下文升级处理,并且 session 和 aggregate routing 的结果得到报告时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Summary
Add a cascade routing mode: sessions run on a user-chosen primary model (typically a local/self-hosted or inexpensive model) and automatically escalate to a user-chosen frontier model when deterministic verification fails. Users pick BOTH models — this is a first-class routing option, not a hardcoded config.
Motivation
- Local/inexpensive models complete a meaningful share of data-engineering tasks correctly; the problem is knowing which ones. altimate-code's deterministic verifier stack (build success, data tests, schema-contract checks) can answer that without ground truth.
- A verifier-gated cascade delivers near-frontier effective quality at a fraction of the cost — while making the cost/latency trade visible to the user instead of implicit.
- Human time matters more than tokens: the design must escalate FAST (default: one primary attempt → gate → escalate) and route around the primary entirely for task classes it historically fails.
Approach
- Model selection is user-driven:
primaryandescalationare both picked by the user (any configured provider/model pair) via config, flags, and interactive selection. No baked-in model names. - Escalation trigger is deterministic verification, not model self-assessment: the gate runs the same checks in both lanes.
- Fast defaults: 1 primary attempt in interactive sessions; batch/unattended mode may raise attempts.
- Measurement built in from day one: per-session routing summary + aggregate reporting (escalation rate, gate outcomes, latency and cost per lane).
- Learn from mistakes: outcomes feed a local routing memory so task classes that consistently escalate skip the primary next time.
Subtasks
- #1138 Cascade routing core: two-model selection + escalate-on-gate-failure flow
- #1135 Verification gate: pluggable deterministic checks as the escalation trigger
- #1136 Measurement: routing telemetry, per-session summary, aggregate report
- #1137 Adaptive routing memory: learn task-class priors from outcomes
Non-goals (v1)
- No learned/ML router — priors are simple aggregates.
- No mid-task model switching within a single attempt; escalation restarts the task on the escalation model with context.
Prior art & evidence (researched 2026-08-24)
A four-lane research pass (papers, vendor first-party writing, harness survey, verifier-gating literature) backs this design. Full 18-system prior-art map lives in the internal brief; key facts:
Novelty claim, stated precisely: execution-based candidate selection is mature (CodeT, AlphaCode, CHASE-SQL), and pre-generation model routing is mature (FrugalGPT, RouteLLM, Cursor Router, GPT-5's router, OpenRouter Auto). But a survey of 10 shipped harnesses/routers (Aider, Cline, Roo-Code, Cursor, OpenRouter, LiteLLM, Copilot, Continue.dev, RA-Aid, gptme/goose) found zero that escalate across model tiers on a deterministic execution signal (build/tests/contract). Every shipped "smart" router classifies before generation. Verifier-gated cross-tier escalation is unshipped territory.
Numbers that justify the pattern:
- FrugalGPT: try-then-escalate cascade matches best-single-LLM accuracy at up to 98% cost reduction (arxiv.org/abs/2305.05176)
- RouteLLM: >2× cost cut, no quality loss; only 13.4% of queries need the strong model for half the quality gap (arxiv.org/abs/2406.18665)
- CodeT: execution-agreement selection lifts HumanEval pass@1 47.0%→65.8% (arxiv.org/abs/2207.10397)
- AlphaCode: execution-test filtering rejects ~99% of samples and is what makes cheap over-sampling viable at all (deepmind AlphaCode paper)
- Query and Conquer: execution-guided self-consistency helps weakest models most (+15.9pts for an 8B vs +3.5pts for a 70B) — directly supports gating a cheap/local primary (arxiv.org/html/2503.24364)
- Cursor Router A/B: 60% cost saving vs always-frontier (cursor.com/blog/router) — but gated on a classifier, not execution
- Anthropic (Building Effective Agents; Agent SDK subagents) and OpenAI (GPT-5 router, model-selection guide) both document tiered-model routing as the recommended pattern; both trigger pre-generation
Evidence-driven design adjustments (applied to subtasks):
- Primary-model attempts default: NOT a strict single shot. Execution-checked multi-sample (2-3 attempts, agreement-gated) disproportionately benefits small/local primaries (Query and Conquer) — use n=2-3 in unattended mode; n=1 remains the interactive default for latency.
- Gate quality is an ongoing investment, not a checkbox: CHASE-SQL's execution selector still leaves ~40-70% of oracle headroom, and TestPrune shows better tests add +2.4 to +12.9pts on an already-gated pipeline. The gate interface must make adding/refining checks cheap.
- Manual model-pair selection is the right v1 (matches Aider/Cline/Continue UX norms); design the config so a learned pre-classifier can be layered later without breaking it — but do NOT ship a learned router early: RouteLLM shows routers are near-random when undertrained, which is exactly the trap for our low-volume early data.
Risks from the literature: router overconfidence on thin data (RouteLLM); verifier gaming by weak self-generated checks (CodeT caveats) — our gate uses project-declared tests, not model-generated ones, which sidesteps the worst of this; latency stacking on serial escalation (mitigated by fast-fail gates and the adaptive skip-primary memory).
- 主要语言
- TypeScript
- 星标
- 805
- 派生
- 122
- 平均合并
- 2 天 6 小时
- 30 天内合并 PR
- 67
环境准备
在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
AltimateAI/altimate-code 的其他 Issue
-
test: MCP tests fail when the developer's ~/.claude.json has MCP servers (HOME is not sandboxed)未关闭
难度 2/5 1-3 小时 新手友好度 78/100
AltimateAI/altimate-code#1386 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
AltimateAI/altimate-code#1384 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
AltimateAI/altimate-code#1323 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
AltimateAI/altimate-code#1288 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 92/100
AltimateAI/altimate-code#1285 ·
维护者通常 1 天内回复
查看 AltimateAI/altimate-code 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
Effect-TS/effect#8881 · 1 条评论 ·
维护者通常 1 天内回复
-
Discover carries headerEdges that nothing reads since #1914 moved E0507/E0517 to the compiler graph未关闭tech-debt
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复
-
mail processing verified
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 7 天内回复
-
check:passed streams:add
难度 1/5 1 小时以内 新手友好度 86/100
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 74/100
OpenSlides/OpenSlides#7180 ·