Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Automation runs misdescribe delivery and act like chat turns

Abierto
#2,014 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Necesita aclaración
Estado de actividad
Activo
Stack tecnológico
typescript
Área
backend

Línea de trabajo

Start by reading packages/junior/src/chat/prompt.ts and packages/junior/src/chat/task-input.ts, then trace delivery from providers/slack/turn.ts, especially how dispatch outcomes select destinations. Review chat/README.md for current silence semantics. Done requires an agreed and tested automation-specific prompt and delivery contract covering unattended work, identity, conditions, silence, credentials, and deliverable-oriented output.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Automation runs (scheduled automations, event automations, watches) get the same system prompt as interactive chat, plus a dispatch wrapper that misdescribes the run. In production this leads to output that makes no sense to the person receiving it: reminders delivered as failure notes, unattended runs asking questions nobody can answer, and status reports where the deliverable should be. Evidence comes from Sentry spans for prod 0.235.0 (prompt and output text is recorded for public conversations only) and from code on main @ ffcd312.

Reported via a failed DM reminder. David Cramer broadened the scope to the general prompt wrapper around automations.

What the model sees today
  1. The static system prompt shared with interactive chat. It includes interactive rules: "Ask the user only for missing access, approval, or a decision…", "Assistant text is delivered only into the active conversation or thread", and "the actor is the person asking now" (packages/junior/src/chat/prompt.ts).
  2. A <dispatch> block (buildDispatchSection): "the runtime delivers the final answer to the destination", actor.name: scheduler|junior, the task's destination.channel_id, and schedule metadata.
  3. <runtime> with slack.conversation.type for the task destination.
  4. <current-instruction author_id="scheduler"> (or "junior") wrapping renderTaskInput (packages/junior/src/chat/task-input.ts). The message-outcome contract is "Briefly report what you did or what is needed next."
Concerns, with production evidence

1. The prompt names the wrong delivery target (high confidence)

  • buildDispatchSection renders the task destination. But providers/slack/turn.ts (~L822) posts to dispatch.outcomes[].destination whenever outcomes exist. So for a task_creator outcome, the prompt describes the creation channel while the reply goes to the creator's DM.
  • dispatch_a60c7cfc0df1e80b0a47e5df66bad14d (2026-09-22): the instruction was "Remind to revisit …". The prompt said public_channel with the creation channel ID. The chat.postMessage span shows the post went to a D… DM. The output was third person (" — reminder to …") and had no mention, so the recipient wasn't pinged.
  • dispatch_55f35d79c2eef338ae8657cd358f0aaa (2026-10-02, private, so text is redacted): the run called userLookup once and posted to a DM. The delivered text was "I found , but this task can only reply in its current conversation, so I couldn't DM him. Please send: …". The reminder reached the right person, framed as a failure.

2. Unattended runs ask for permission or questions (high confidence)

  • dispatch_3f68e08acfff6ad09eb93604c3ae492f (2026-09-30, daily scheduled sweep): the instruction says to run a skill and log to Notion. The output: "Today's sweep is on hold… Can I search the … database in Notion? The search is read-only… After you confirm, I'll…". The daily check didn't run.
  • dispatch_461b6009706a6366675081cccd1d47fb (2026-09-26, scheduled PR review): the instruction says "Stay silent for … blockers" and to post only when a PR is ready. The output posted "may i replace PR getsentry/sentry-mcp#1253's description…?"
  • Event automation dispatch_932385b71b0723b53ae0401db0c76508 (2026-09-30): the instruction says to act only when two conditions hold, otherwise "do nothing". The output: "Should I open the draft PR…? Please first confirm that . I can't check … because access isn't authorized in this run." It had already pushed a branch before checking the condition.

3. The actor is scheduler / junior, not the creator (high confidence)

  • There is no <actor> block, and author_id is a system name, so "me", "you", and "whose DM is this" can't be resolved. #2009 (on main, not yet released) adds a Created by: line to the task text. The actor block and author_id still say scheduler / junior. Related: #1394.

4. The reply contract asks for a process report, not the deliverable (high confidence)

  • "Briefly report what you did or what is needed next" pushes reminders and digests toward meta-narration. The "Please send: …" output above is a "what is needed next" answer.
  • Silent-outcome runs still write reports that narrate compliance with their instructions. Examples: dispatch_28e64f9f2ec91641273cc0322da4e46a and dispatch_311463700e053dbeef579e64dd3c46fc, warden PR reviews (2026-09-30): "I posted a comment review… (no approval, no merge, nothing in Slack)" and "Author check: … not Junior, so it qualified". These weren't delivered, but they show the same tendency.

5. Silence and delivery semantics are pushed onto users (high confidence)

  • chat/README.md says "Task input never asks the model to emit a silence marker". Conditional silence only works through the system prompt's [[NO_REPLY]] rule, so users hand-write delivery rules into instructions: "reply in this Slack conversation tagging @… / otherwise reply with exactly [[NO_REPLY]]", "Only post to #channel at channel top level when…", "Stay silent for…", and "Do not post or reply in Slack" (on a task whose outcome is already silent).
  • In the last 100 public runs of each type, 83/100 event-automation runs and 16/47 scheduled runs ended in [[NO_REPLY]]. It mostly works, but only when users know the incantation.

6. Runs discover missing provider credentials mid-run (medium confidence, one example)

  • The prompt doesn't state which provider credentials a run has (creator vs system), so the model hits the gap partway through and improvises. Example: " access isn't authorized in this run" in the event-automation case above.
User-expected intent
  • Delivery: the person creating an automation expects the stored outcome to decide where the message goes, and the message to read naturally there. "Remind me in a DM" means a DM that is the reminder, addressed to them, not a note about delivery.
  • Unattended: the instructions are the creator's standing approval within their scope. Nobody is watching the run. It should do the work, or stop cleanly and not ask.
  • Conditions first: "only act when X" means check X before any side effect. If X can't be verified, take no action.
  • Silence: "stay silent unless…" should be a supported behavior. Users shouldn't need to know about [[NO_REPLY]].
  • Identity: "me" in an instruction refers to the creator, and mentions should ping the right person.
  • Output: the visible message is the deliverable (reminder, digest, alert, result), not a narration of the process or of the instructions it followed.
Unknowns
  • How often the delivery-target mismatch occurs across all automations. Prompt text for private conversations is redacted, so DM cases can only be confirmed from span metadata (tool calls, post channel, output length).
  • The exact stored instruction for the 2026-10-02 DM reminder.

via David Cramer.

--

View Junior Session [Sentry]

Lenguaje dominante
TypeScript
Estrellas
367
Forks
41
Merge medio
6 h 9 min
PR fusionados (30 d)
188

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de getsentry/junior

Todos los issues de getsentry/junior

Issues similares

Más issues de TypeScript

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.