LocalSessionManager session workers don't cancel in-flight tool calls on client disconnect (follow-up to #857)
Maintainers usually reply within 3 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- rust
- Domain
- backend-api-design, networking
Research direction
Read transport/streamable_http_server/session/local.rs alongside transport/streamable_http_server/tower.rs, especially serve_negotiated_request_directly, to compare the existing cancellation lifecycle with session-worker dispatch. Reproduce the issue with the provided LocalSessionManager example; done means a client disconnect cancels the in-flight tool call on the session-manager path, as required by the issue's expected behavior.
Written by the indexing model from the issue text.
Description
Summary
#857 reported that a client's TCP disconnect doesn't cancel an in-flight #[tool] future on the streamable-HTTP transport, and was closed after a fix landed in transport/streamable_http_server/tower.rs's serve_negotiated_request_directly (the stateless direct-dispatch path), which explicitly cites #857 in a comment and uses a CancellationToken + drop_guard() tied to the request's lifecycle.
That fix does not appear to cover the session-manager dispatch path (LocalSessionManager, in transport/streamable_http_server/session/local.rs). That file dispatches each session through WorkerTransport::spawn(worker) (lines 68, 175) with no CancellationToken/drop_guard wiring anywhere in it — only catch_cancellation_notification, which handles the explicit notifications/cancelled message, not a raw transport disconnect. I confirmed this by reproducing the original bug against current rmcp 3.3.0, using a server built with LocalSessionManager::default() (likely a common setup, since it's the session manager shown in-tree).
This also appears to be a spec compliance gap, independent of the original issue: the 2026-07-28 Streamable HTTP cancellation spec states the server MUST treat a client disconnect as cancellation of that request for this transport.
Versions
rmcp = "3" (resolved 3.3.0), axum = "0.8" (resolved 0.8.9), tokio = "1" (resolved 1.53.1), via transport-streamable-http-server.
Reproduction
Cargo.toml:
[dependencies]
rmcp = { version = "3", features = ["server", "macros", "schemars", "transport-streamable-http-server"] }
axum = "0.8"
serde = { version = "1", features = ["derive"] }
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
src/main.rs:
use rmcp::{
handler::server::wrapper::Parameters,
schemars, tool, tool_router,
transport::streamable_http_server::{
StreamableHttpServerConfig, StreamableHttpService, session::local::LocalSessionManager,
},
};
#[derive(Debug, serde::Deserialize, schemars::JsonSchema)]
struct SleepRequest {
ms: u64,
}
#[derive(Clone)]
struct UnstableServer;
#[tool_router(server_handler)]
impl UnstableServer {
#[tool(description = "Sleep for the given number of milliseconds, then respond")]
async fn sleep(&self, Parameters(SleepRequest { ms }): Parameters<SleepRequest>) -> String {
println!("sleep: starting {ms}ms");
tokio::time::sleep(std::time::Duration::from_millis(ms)).await;
println!("sleep: finished {ms}ms");
format!("slept {ms}ms")
}
}
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let service = StreamableHttpService::new(
|| Ok(UnstableServer),
LocalSessionManager::default().into(),
StreamableHttpServerConfig::default(),
);
let router = axum::Router::new().nest_service("/mcp", service);
let listener = tokio::net::TcpListener::bind("127.0.0.1:8002").await?;
axum::serve(listener, router).await?;
Ok(())
}
Run it, then from another terminal, send a tools/call for sleep with a 5-second delay, but force the client to give up after 1 second:
curl -s --max-time 1 -X POST http://127.0.0.1:8002/mcp \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-H 'MCP-Protocol-Version: 2026-07-28' -H 'Mcp-Method: tools/call' -H 'Mcp-Name: sleep' \
--data-raw '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"sleep","arguments":{"ms":5000},"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{},"io.modelcontextprotocol/clientInfo":{"name":"repro","version":"0.1.0"}}}}'
Expected
curl exits after ~1s (--max-time fired); the server log stops after sleep: starting 5000ms — the handler is cancelled along with the dropped connection.
Actual
curl does exit at ~1s, but the server log still prints sleep: finished 5000ms a full 5 seconds after the request started.
I additionally monitored the server process with lsof -p <pid> -a -i tcp at 0.5s intervals against this exact repro: the ESTABLISHED connection from the client is present at the 0.5s and 1.0s samples and gone by the 1.5s sample — the TCP connection genuinely closes within half a second of the client's disconnect, roughly 3.5 seconds before the tool handler prints finished. So this isn't a case of the disconnect being missed at the transport level; the connection closes correctly and promptly, but the tool invocation itself isn't wired to anything that would stop it.
Suggested fix direction
Mirror the pattern already in tower.rs's serve_negotiated_request_directly: a per-request CancellationToken with a drop_guard() tied to the connection/response lifecycle, applied to the session worker's dispatch in session/local.rs instead of (or in addition to) the stateless direct path. Happy to attempt a PR along these lines if that shape sounds right — flagging here first per CONTRIBUTE.MD's "discuss before a PR" guidance, and since session/local.rs's session-affinity model may need a different shape than the stateless path's.
- Dominant language
- Rust
- Stars
- 4k
- Forks
- 654
- Avg merge
- 4d 2h
- Merged PRs (30d)
- 40
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from modelcontextprotocol/rust-sdk
-
P3 question T-documentation T-enhancement
Difficulty 1/5 Under an hour Newbie friendliness 86/100
modelcontextprotocol/rust-sdk#1155 ·
Maintainers usually reply within 3 days
-
streamable-http server: client responses to server-initiated requests (sampling/elicitation/roots) are 202-accepted and silently discarded under the 2026-07-28 protocol — the pending request hangs foreverPossibly taken A pull request linked to this issue is open or already merged. Openbug P1 ready for work T-transport
Difficulty 5/5 Over a week Newbie friendliness 35/100
modelcontextprotocol/rust-sdk#1321 ·
Maintainers usually reply within 3 days
-
Bound pre-lifecycle bootstrap attempts in server initializationPossibly taken @DaleSeo claimed this 4 days ago. Openenhancement P2 T-service T-transport
modelcontextprotocol/rust-sdk#1315 · 1 assignee ·
Maintainers usually reply within 3 days
-
ProgressDispatcher: a slow progress subscriber blocks subscribe() and delivery for other tokensOpenbug P1 ready for work T-handler
Difficulty 4/5 3-5 days Newbie friendliness 25/100
modelcontextprotocol/rust-sdk#1312 ·
Maintainers usually reply within 3 days
-
transport::stdio() runs every read and write on tokio's blocking pool, which caps stdio throughputPossibly taken A pull request linked to this issue is open or already merged. Openenhancement P2 T-transport
Difficulty 4/5 3-5 days Newbie friendliness 58/100
modelcontextprotocol/rust-sdk#1301 · 1 comment ·
Maintainers usually reply within 3 days
All issues in modelcontextprotocol/rust-sdk
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 5 days
-
state:triage-needed
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
ktuner keeps a stale ledger path and can never restore that entryPossibly taken @Frun1na claimed this today. Opencomponent:ktuner
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
agentic-os-org/ANOLISA#6483 · 1 comment ·
Maintainers usually reply within 1 day
-
[Resource]: snapbackOpenresource-submission validation-passed
Difficulty 1/5 Under an hour Newbie friendliness 85/100
hesreallyhim/awesome-claude-code#3095 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day