Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Cascade routing mode: user-selected primary model with verifier-gated frontier escalation

未关闭
#1,134 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
typescript

调研方向

从子任务 #1138、#1135、#1136 和 #1137 开始,然后跟踪现有的 config、flags 和交互式模型选择入口。检查 verifier stack 如何报告 build、data-test 和 schema-contract 的结果。当用户可以选择两个模型、失败的验证会携带上下文升级处理,并且 session 和 aggregate routing 的结果得到报告时,即表示完成。

由索引模型根据 Issue 内容生成。

描述

Summary

Add a cascade routing mode: sessions run on a user-chosen primary model (typically a local/self-hosted or inexpensive model) and automatically escalate to a user-chosen frontier model when deterministic verification fails. Users pick BOTH models — this is a first-class routing option, not a hardcoded config.

Motivation

  • Local/inexpensive models complete a meaningful share of data-engineering tasks correctly; the problem is knowing which ones. altimate-code's deterministic verifier stack (build success, data tests, schema-contract checks) can answer that without ground truth.
  • A verifier-gated cascade delivers near-frontier effective quality at a fraction of the cost — while making the cost/latency trade visible to the user instead of implicit.
  • Human time matters more than tokens: the design must escalate FAST (default: one primary attempt → gate → escalate) and route around the primary entirely for task classes it historically fails.

Approach

  1. Model selection is user-driven: primary and escalation are both picked by the user (any configured provider/model pair) via config, flags, and interactive selection. No baked-in model names.
  2. Escalation trigger is deterministic verification, not model self-assessment: the gate runs the same checks in both lanes.
  3. Fast defaults: 1 primary attempt in interactive sessions; batch/unattended mode may raise attempts.
  4. Measurement built in from day one: per-session routing summary + aggregate reporting (escalation rate, gate outcomes, latency and cost per lane).
  5. Learn from mistakes: outcomes feed a local routing memory so task classes that consistently escalate skip the primary next time.

Subtasks

  • #1138 Cascade routing core: two-model selection + escalate-on-gate-failure flow
  • #1135 Verification gate: pluggable deterministic checks as the escalation trigger
  • #1136 Measurement: routing telemetry, per-session summary, aggregate report
  • #1137 Adaptive routing memory: learn task-class priors from outcomes

Non-goals (v1)

  • No learned/ML router — priors are simple aggregates.
  • No mid-task model switching within a single attempt; escalation restarts the task on the escalation model with context.

Prior art & evidence (researched 2026-08-24)

A four-lane research pass (papers, vendor first-party writing, harness survey, verifier-gating literature) backs this design. Full 18-system prior-art map lives in the internal brief; key facts:

Novelty claim, stated precisely: execution-based candidate selection is mature (CodeT, AlphaCode, CHASE-SQL), and pre-generation model routing is mature (FrugalGPT, RouteLLM, Cursor Router, GPT-5's router, OpenRouter Auto). But a survey of 10 shipped harnesses/routers (Aider, Cline, Roo-Code, Cursor, OpenRouter, LiteLLM, Copilot, Continue.dev, RA-Aid, gptme/goose) found zero that escalate across model tiers on a deterministic execution signal (build/tests/contract). Every shipped "smart" router classifies before generation. Verifier-gated cross-tier escalation is unshipped territory.

Numbers that justify the pattern:

  • FrugalGPT: try-then-escalate cascade matches best-single-LLM accuracy at up to 98% cost reduction (arxiv.org/abs/2305.05176)
  • RouteLLM: >2× cost cut, no quality loss; only 13.4% of queries need the strong model for half the quality gap (arxiv.org/abs/2406.18665)
  • CodeT: execution-agreement selection lifts HumanEval pass@1 47.0%→65.8% (arxiv.org/abs/2207.10397)
  • AlphaCode: execution-test filtering rejects ~99% of samples and is what makes cheap over-sampling viable at all (deepmind AlphaCode paper)
  • Query and Conquer: execution-guided self-consistency helps weakest models most (+15.9pts for an 8B vs +3.5pts for a 70B) — directly supports gating a cheap/local primary (arxiv.org/html/2503.24364)
  • Cursor Router A/B: 60% cost saving vs always-frontier (cursor.com/blog/router) — but gated on a classifier, not execution
  • Anthropic (Building Effective Agents; Agent SDK subagents) and OpenAI (GPT-5 router, model-selection guide) both document tiered-model routing as the recommended pattern; both trigger pre-generation

Evidence-driven design adjustments (applied to subtasks):

  1. Primary-model attempts default: NOT a strict single shot. Execution-checked multi-sample (2-3 attempts, agreement-gated) disproportionately benefits small/local primaries (Query and Conquer) — use n=2-3 in unattended mode; n=1 remains the interactive default for latency.
  2. Gate quality is an ongoing investment, not a checkbox: CHASE-SQL's execution selector still leaves ~40-70% of oracle headroom, and TestPrune shows better tests add +2.4 to +12.9pts on an already-gated pipeline. The gate interface must make adding/refining checks cheap.
  3. Manual model-pair selection is the right v1 (matches Aider/Cline/Continue UX norms); design the config so a learned pre-classifier can be layered later without breaking it — but do NOT ship a learned router early: RouteLLM shows routers are near-random when undertrained, which is exactly the trap for our low-volume early data.

Risks from the literature: router overconfidence on thin data (RouteLLM); verifier gaming by weak self-generated checks (CodeT caveats) — our gate uses project-declared tests, not model-generated ones, which sidesteps the worst of this; latency stacking on serial escalation (mitigated by fast-fail gates and the adaptive skip-primary memory).

主要语言
TypeScript
星标
805
派生
122
平均合并
2 天 6 小时
30 天内合并 PR
67

环境准备

在 Codespaces 中打开

在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

AltimateAI/altimate-code 的其他 Issue

查看 AltimateAI/altimate-code 的全部 Issue

相似的 Issue

更多 TypeScript Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。