Scope the builder Pre-Execution Protocol to task shape
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 66/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- sql, typescript
- Ambito
- cli, developer-experience
Direzione di ricerca
Inizia da packages/opencode/src/altimate/prompts/builder.txt e analizza il precedente esistente di SessionTermination.completionInstruction per l’iniezione del prompt specifica della modalità di esecuzione. Traccia come vengono determinati headless mode, il builder agent e la presenza di un progetto dbt. Il lavoro è completo quando il Pre-Execution Protocol viene omesso solo per le esecuzioni di headless builder senza un progetto dbt e rimane presente in tutti gli altri casi, inclusi i workspace non classificati.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
packages/opencode/src/altimate/prompts/builder.txt carries a ## Pre-Execution Protocol section that makes sql_analyze + altimate_core_validate mandatory before every sql_execute. builder is a PRIMARY agent, so that section reaches every builder surface at once: dbt authoring, interactive chat, and headless question-answering runs.
An internal pre-registered paired prompt ablation (540 trials on a public data-question benchmark, one binary across both arms, per-trial system prompts verified by sha256) measured what the section costs on the question-answering surface:
| macro Pass@1 | |
|---|---|
| control | 0.6667 |
| treatment (protocol removed, among other changes) | 0.6807 |
| delta | +0.0140 |
- query-blocked sign-flip permutation, 20,000 resamples: p = 0.7358
- cluster-bootstrap 95% CI: [-0.0400, +0.0674]
That is a null on score. The point estimate is positive and the interval comfortably contains zero — it is not an improvement and must not be reported as one.
What did move:
| control | treatment | delta | |
|---|---|---|---|
| wall clock per trial (restricted mean) | 440.9s | 319.4s | -27.6% |
| model turns per trial | 11.9 | 8.6 | -27.7% |
| generation seconds per trial | 187.6 | 127.2 | -32.2% |
altimate_core_validate calls (arm total) |
1,476 | 0 | -1,476 |
sql_analyze calls (arm total) |
1,329 | 0 | -1,329 |
sql_execute calls (arm total) |
1,716 | 2,554 | +49% |
| trials hitting the 900s timeout | 27 | 16 | -11 |
The 2,805 ritual tool calls going to zero is the number directly attributable to this text. The ritual is prompt-ordered: deleting the order deletes it completely, not partially. The freed budget went into actual querying (sql_execute +49%).
Why not just delete it
Two reasons, both from the experiment's own caveats:
- The latency win is not attributable to this section alone. The treatment arm bundled five coupled changes; the experiment explicitly declined to attribute and a factorial was not run.
- The measurement covers data questions only. dbt authoring and interactive chat are unmeasured builder surfaces where a pre-execution discipline may genuinely earn its place. Deleting on this evidence would over-generalise from one workload.
Proposal
Scope the section to task shape rather than delete it, following the precedent already in the tree: SessionTermination.completionInstruction moved a run-mode-only instruction out of builder.txt and injects it only when headless AND the agent is builder.
Drop the protocol only in the cell that was measured (headless + builder + no dbt project in the workspace) and keep it everywhere else, including whenever the workspace cannot be classified. The cost of keeping it is 27% latency on one workload; the cost of wrongly dropping it is unmeasured.
Out of scope
main also carries a ## Finish Protocol section (shipped in #1171) — a second mandatory ritual in the same family, added after the binary the ablation measured was built. No measurement covers it and this issue does not propose touching it.
A follow-up measurement on dbt tasks is needed before the gate could be widened, or the section deleted outright.
- Lingua principale
- TypeScript
- Stelle
- 813
- Fork
- 134
- Merge medio
- 1g 19h
- PR unite (30g)
- 64
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di AltimateAI/altimate-code
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
AltimateAI/altimate-code#1359 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
AltimateAI/altimate-code#1323 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
AltimateAI/altimate-code#1288 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
AltimateAI/altimate-code#1285 ·
I maintainer di solito rispondono entro 1 giorno
-
privacy: Altimate Base consent dialog no longer discloses persistent per-installation identifierAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
AltimateAI/altimate-code#1284 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di AltimateAI/altimate-code
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
rohitg00/agentmemory#1428 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
boxlite-ai/boxlite#1729 ·
I maintainer di solito rispondono entro 1 giorno
-
detectors enhancement good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
SM260845/readme-gen#1 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
angular/angularfire#3774 ·
I maintainer di solito rispondono entro 2 giorni