fix(gateway): prevent agent-initiated self-restart from killing the active session
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python, zsh
- Domain
- backend, devops, testing-qa
Research direction
Start with the raven gateway entry point and trace how an active channel turn handles restart requests and process ownership. Review the recoverability proposal in #174 and reproduce the active-turn scenario with the configured shell wrapper. Done means a restart is rejected or supervisor-managed without losing the final status, with lifecycle tests covering the case and no forced-stop fallback by default.
Written by the indexing model from the issue text.
Description
Summary
When an agent with host exec access patches Raven and then tries to restart the running gateway from inside that same gateway, it can terminate the process that is executing the current turn. The active channel session is interrupted before the agent can verify the restart or report the outcome.
In the observed path, the agent first requested graceful termination of the gateway wrapper, waited for it to exit, then escalated to forced termination for both the wrapper and gateway child. The process did come back through its external parent, but that recovery was incidental rather than a Raven-managed restart contract. The forced shutdown also emitted Python resource-tracker warnings about leaked semaphores.
This is not specific to Telegram reactions; that feature merely led the agent to apply a local patch that required a restart.
Steps to reproduce
- Start
raven gatewaythrough a shell wrapper, with an agent channel enabled and theexectool allowed on the host. - Ask the agent to make a source/configuration change that requires the running gateway to reload.
- Let the agent attempt to restart the gateway by discovering its wrapper/child PIDs, requesting termination, and escalating to forced termination when graceful shutdown does not finish quickly.
- Observe that the process running the current agent turn is terminated; the user receives no completion/result for that turn.
Expected behavior
Raven should provide a safe, explicit restart path rather than leaving an in-process agent to stop its own serving gateway. A restart request should either:
- be rejected/deferred while processing the current turn, with an actionable instruction for the operator; or
- be handed to a managed supervisor/control plane that can drain or persist work, restart the gateway, and report the result after recovery.
The default agent tool surface should not make self-termination the easiest implementation of a requested restart.
Actual behavior
The agent terminated its own wrapper and gateway child. The current turn stopped mid-task, and the channel was unavailable until a parent process started another gateway instance. Forced termination also produced shutdown cleanup warnings.
Relevant sanitized log excerpt:
Tool call: execute a graceful-stop request for the old wrapper
... the old gateway wrapper did not exit after 10 seconds
...
Tool call: force-stop the gateway child and wrapper
shell: terminated .../raven-gateway-wrapper
...
UserWarning: resource_tracker: There appear to be leaked semaphore objects to clean up at shutdown
Environment
- OS: macOS arm64
- Shell: zsh
- Python: 3.12.13
- Installation: uv tool install
- Channel: Telegram long polling
- Raven: local installed package; issue reproduced on a July 2026 build while applying a local extension
- Gateway launched by a shell wrapper that creates a temporary runtime config
Suggested direction
Consider a first-class gateway lifecycle interface with a clear ownership boundary, for example:
- A
restart_requestedcontrol action that cannot directly signal the current serving process. - Supervisor integration (or an operator-only command) that owns stop/start and health verification.
- A safe default that reports “restart required” and preserves the current turn when no supervisor is configured.
- Lifecycle tests covering a restart requested from an active channel turn, including no lost final status and no forced-stop fallback by default.
This may relate to the broader recoverability proposal in #174, but the immediate problem exists even before durable inbox/outbox work: an agent should not be able to accidentally sever the conversation that is asking it to perform the change.
- Dominant language
- Python
- Stars
- 4.1k
- Forks
- 94
- Avg merge
- 10h 2m
- Merged PRs (30d)
- 376
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EverMind-AI/Raven
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#798 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
EverMind-AI/Raven#797 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
EverMind-AI/Raven#640 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#479 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 75/100
EverMind-AI/Raven#474 · 2 comments ·
Maintainers usually reply within 1 day
All issues in EverMind-AI/Raven
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
letsencrypt/cp-cps#353 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
DOI-USGS/pywatershed#421 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
python-pillow/Pillow#10087 · 1 comment ·
Maintainers usually reply within 1 day