fix(playbook): the generation budget is shorter than a clean generation takes
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
Research direction
Start with the reproduced PlaybookGenerator.generate call and inspect CreatePlaybookTool.timeout_seconds, the generation prompt, agent profiles, and inline skill docs. Compare the logged round count with the 338-second elapsed time to identify whether prompt construction or the timeout is responsible. Done means explaining the delay and determining an evidence-based timeout or round-bound change for both creation entries.
Written by the indexing model from the issue text.
Description
Summary
CreatePlaybookTool.timeout_seconds is 180.0, with the comment "Generation is
a few LLM calls (draft plus repair rounds)". Measured on a configured machine, a
generation that needed no repair rounds at all took 338 seconds -- so the
budget is short by most of a clean run, and creating a playbook from a
conversation cannot succeed there.
Nothing is wrong with the machine's model setup, which is what makes this worth
filing rather than dismissing.
Measurement
Host: an ordinary hosted provider reached over the public internet, with a
fast general-purpose model. The exact vendor is not the point -- the checks
below show the endpoint and the provider are both healthy.
generation input: "Fetch the open issues for a repository, then write a
one-paragraph weekly summary. The repository name changes
per run."
playbook weekly-issue-summary generated in 1 round(s)
elapsed: 338.2s
nodes: ['fetch-issues', 'weekly-summary']
The result is good -- two sensible nodes and three review notes. It simply
arrives after the budget has expired.
Ruled out, in this order, so the number is the remaining explanation:
| Checked | Result |
|---|---|
| Endpoint reachable | provider's model listing -> 200 in 2.0s |
| Provider usable from this build | one-message chat() round trip -> OK in 2.4s |
| Model configured | resolved through make_provider, a current hosted model |
| Repair rounds inflating it | none -- the log says generated in 1 round(s) |
Why it matters beyond the tool
The budget is the only bound on creation, and it is declared once and read by
more than one entry, so raising or lowering it moves every creation path at
once. Today the conversational entry (create_playbook) is the one that hits
it; any other entry that binds the same composer inherits the same ceiling.
The failure is also quiet in the worst way for a user: the wait is spent, the
generation is thrown away at the deadline, and the model composed a perfectly
good playbook that nobody ever sees.
What to decide
- Raise the number. Simplest, and the measurement says roughly 2x is the
floor for one clean round on this configuration -- but a number picked from
one host's timing is a guess for the next one. - Bound the rounds instead of the wall clock, or bound both. The comment's
own model of the cost ("a draft plus repair rounds") is round-shaped; a
wall-clock ceiling converts a slow endpoint into a lost generation even when
the work was progressing normally. - Find out why one round costs 338s first. 338s for a single round with no
repairs is far enough from the intuition behind "a few LLM calls" that the
prompt is worth looking at before the timeout is. The generation prompt
carries the system prompt, the agent profiles and inline skill docs; if the
skill material is being inlined more broadly than intended, the fix is there
and the budget never needed changing.
(3) is the one worth doing first -- it may remove the need for (1) entirely, and
if it does not, it gives (1) a real number to pick instead of a doubled guess.
Reproducing
Build the composer the way both creation entries do and time one generate:
gen = PlaybookGenerator(make_provider(cfg), None,
lambda: agent_profiles_from_registry(reg),
live_inventory(cfg.tools.mcp_servers),
model=cfg.playbooks.model)
await gen.generate("<any two-step workflow>")
The log line playbook <name> generated in N round(s) reports the round count,
which is what separates "slow model" from "many repairs".
- Dominant language
- Python
- Stars
- 4.1k
- Forks
- 94
- Avg merge
- 9h 54m
- Merged PRs (30d)
- 370
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EverMind-AI/Raven
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#798 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
EverMind-AI/Raven#797 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
EverMind-AI/Raven#640 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#479 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 75/100
EverMind-AI/Raven#474 · 2 comments ·
Maintainers usually reply within 1 day
All issues in EverMind-AI/Raven
Similar issues
-
customer-reported
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Azure/azure-cli#34150 · 1 comment ·
Maintainers usually reply within 1 day
-
community-request
Difficulty 1/5 Under an hour Newbie friendliness 95/100
NVIDIA-NeMo/Curator#2464 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
WeblateOrg/translation-finder#1099 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
trezor/trezor-firmware#7997 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day