Automation runs misdescribe delivery and act like chat turns
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- typescript
- Ambito
- backend
Direzione di ricerca
Start by reading packages/junior/src/chat/prompt.ts and packages/junior/src/chat/task-input.ts, then trace delivery from providers/slack/turn.ts, especially how dispatch outcomes select destinations. Review chat/README.md for current silence semantics. Done requires an agreed and tested automation-specific prompt and delivery contract covering unattended work, identity, conditions, silence, credentials, and deliverable-oriented output.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Automation runs (scheduled automations, event automations, watches) get the same system prompt as interactive chat, plus a dispatch wrapper that misdescribes the run. In production this leads to output that makes no sense to the person receiving it: reminders delivered as failure notes, unattended runs asking questions nobody can answer, and status reports where the deliverable should be. Evidence comes from Sentry spans for prod 0.235.0 (prompt and output text is recorded for public conversations only) and from code on main @ ffcd312.
Reported via a failed DM reminder. David Cramer broadened the scope to the general prompt wrapper around automations.
What the model sees today
- The static system prompt shared with interactive chat. It includes interactive rules: "Ask the user only for missing access, approval, or a decision…", "Assistant text is delivered only into the active conversation or thread", and "the actor is the person asking now" (
packages/junior/src/chat/prompt.ts). - A
<dispatch>block (buildDispatchSection): "the runtime delivers the final answer to the destination",actor.name: scheduler|junior, the task'sdestination.channel_id, and schedule metadata. <runtime>withslack.conversation.typefor the task destination.<current-instruction author_id="scheduler">(or"junior") wrappingrenderTaskInput(packages/junior/src/chat/task-input.ts). The message-outcome contract is "Briefly report what you did or what is needed next."
Concerns, with production evidence
1. The prompt names the wrong delivery target (high confidence)
buildDispatchSectionrenders the taskdestination. Butproviders/slack/turn.ts(~L822) posts todispatch.outcomes[].destinationwhenever outcomes exist. So for atask_creatoroutcome, the prompt describes the creation channel while the reply goes to the creator's DM.dispatch_a60c7cfc0df1e80b0a47e5df66bad14d(2026-09-22): the instruction was "Remind to revisit …". The prompt saidpublic_channelwith the creation channel ID. Thechat.postMessagespan shows the post went to aD…DM. The output was third person (" — reminder to …") and had no mention, so the recipient wasn't pinged.dispatch_55f35d79c2eef338ae8657cd358f0aaa(2026-10-02, private, so text is redacted): the run calleduserLookuponce and posted to a DM. The delivered text was "I found , but this task can only reply in its current conversation, so I couldn't DM him. Please send: …". The reminder reached the right person, framed as a failure.
2. Unattended runs ask for permission or questions (high confidence)
dispatch_3f68e08acfff6ad09eb93604c3ae492f(2026-09-30, daily scheduled sweep): the instruction says to run a skill and log to Notion. The output: "Today's sweep is on hold… Can I search the … database in Notion? The search is read-only… After you confirm, I'll…". The daily check didn't run.dispatch_461b6009706a6366675081cccd1d47fb(2026-09-26, scheduled PR review): the instruction says "Stay silent for … blockers" and to post only when a PR is ready. The output posted "may i replace PR getsentry/sentry-mcp#1253's description…?"- Event automation
dispatch_932385b71b0723b53ae0401db0c76508(2026-09-30): the instruction says to act only when two conditions hold, otherwise "do nothing". The output: "Should I open the draft PR…? Please first confirm that . I can't check … because access isn't authorized in this run." It had already pushed a branch before checking the condition.
3. The actor is scheduler / junior, not the creator (high confidence)
- There is no
<actor>block, andauthor_idis a system name, so "me", "you", and "whose DM is this" can't be resolved. #2009 (onmain, not yet released) adds aCreated by:line to the task text. The actor block andauthor_idstill sayscheduler/junior. Related: #1394.
4. The reply contract asks for a process report, not the deliverable (high confidence)
- "Briefly report what you did or what is needed next" pushes reminders and digests toward meta-narration. The "Please send: …" output above is a "what is needed next" answer.
- Silent-outcome runs still write reports that narrate compliance with their instructions. Examples:
dispatch_28e64f9f2ec91641273cc0322da4e46aanddispatch_311463700e053dbeef579e64dd3c46fc, warden PR reviews (2026-09-30): "I posted a comment review… (no approval, no merge, nothing in Slack)" and "Author check: … not Junior, so it qualified". These weren't delivered, but they show the same tendency.
5. Silence and delivery semantics are pushed onto users (high confidence)
chat/README.mdsays "Task input never asks the model to emit a silence marker". Conditional silence only works through the system prompt's[[NO_REPLY]]rule, so users hand-write delivery rules into instructions: "reply in this Slack conversation tagging @… / otherwise reply with exactly [[NO_REPLY]]", "Only post to #channel at channel top level when…", "Stay silent for…", and "Do not post or reply in Slack" (on a task whose outcome is already silent).- In the last 100 public runs of each type, 83/100 event-automation runs and 16/47 scheduled runs ended in
[[NO_REPLY]]. It mostly works, but only when users know the incantation.
6. Runs discover missing provider credentials mid-run (medium confidence, one example)
- The prompt doesn't state which provider credentials a run has (creator vs system), so the model hits the gap partway through and improvises. Example: " access isn't authorized in this run" in the event-automation case above.
User-expected intent
- Delivery: the person creating an automation expects the stored outcome to decide where the message goes, and the message to read naturally there. "Remind me in a DM" means a DM that is the reminder, addressed to them, not a note about delivery.
- Unattended: the instructions are the creator's standing approval within their scope. Nobody is watching the run. It should do the work, or stop cleanly and not ask.
- Conditions first: "only act when X" means check X before any side effect. If X can't be verified, take no action.
- Silence: "stay silent unless…" should be a supported behavior. Users shouldn't need to know about
[[NO_REPLY]]. - Identity: "me" in an instruction refers to the creator, and mentions should ping the right person.
- Output: the visible message is the deliverable (reminder, digest, alert, result), not a narration of the process or of the instructions it followed.
Unknowns
- How often the delivery-target mismatch occurs across all automations. Prompt text for private conversations is redacted, so DM cases can only be confirmed from span metadata (tool calls, post channel, output length).
- The exact stored instruction for the 2026-10-02 DM reminder.
via David Cramer.
--
- Lingua principale
- TypeScript
- Stelle
- 367
- Fork
- 41
- Merge medio
- 6h 9m
- PR unite (30g)
- 188
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di getsentry/junior
-
documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 64/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
getsentry/junior#335 · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di getsentry/junior
Issue simili
-
Mondriaan
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
knaw-huc/textannoviz#709 ·
I maintainer di solito rispondono entro 1 giorno
-
Add: YRF Music NepalApertastreams:add
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno
-
About page content is also published as an unstyled duplicate at /about/about/Forse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
walletbeat/walletbeat#1558 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
hawk-digital-environments/HAWKI#438 ·
I maintainer di solito rispondono entro 1 giorno
-
good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
OktoLabsAI/okto-pulse#114 ·
I maintainer di solito rispondono entro 1 giorno