MCP: single engine mutex + no op timeout lets one stalled hyperd call hang all tool calls
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 静か
- 技術スタック
- rust
調査の方向性
Start in hyperdb-mcp/src/server.rs at with_engine around lines 821 and 1228, then read hyperdb-api/src/connection_builder.rs around line 136 and the existing ConnectionLost reconnect path. Compare the timeout, pooling, fast-path, and watchdog options before choosing a design. Done means a wedged connection no longer indefinitely blocks unrelated calls, long analytics queries remain valid, and a concurrent-call test verifies the behavior.
索引モデルが issue の本文から書いたものです。
説明
Summary
The MCP server serializes every tool call behind a single
Arc<Mutex<Option<Engine>>> (server.rs:821),
and with_engine holds that lock across the entire blocking hyperd operation
(server.rs:1228). With no
per-operation timeout, a single slow or stalled hyperd call blocks all
other tool calls — including lightweight ones like status — for as long as
the slow call runs.
How it surfaced
A user reported status "hanging." Investigation showed status itself is
trivial, but it was queued behind another operation on a long-idle connection
that had gone half-open (laptop sleep / network blip). The immediate trigger —
no TCP keepalive, so a half-open socket blocked for the 2h OS idle default — is
fixed separately (TCP keepalive, ~90s dead-peer detection). This issue is the
amplifier: the global mutex turns one stalled connection into a total stall of
the MCP surface, and the missing timeout removes the only other backstop.
This became more impactful in v0.5.0, where the daemon became
resident-by-default and connections now live indefinitely across suspends.
Why this is filed as a follow-up (not fixed in the keepalive PR)
Keepalive bounds the worst case to ~90s and is a safe, surgical change.
Removing the serialization is an architectural change (concurrency model of
the engine) and deserves its own design pass rather than a rushed patch. Filing
to track it.
Options to evaluate (not a decision)
- Per-operation timeout / cancellation. Wrap blocking
hyperdcalls so a
stalled op can't hold the lock unboundedly. The connection builder already
exposesquery_timeout(connection_builder.rs:136)
— but a blanket query timeout is wrong for HyperDB (legitimate long
analytics queries). Any timeout must target liveness (is the peer
responding) not duration (how long the query runs). - Connection pool instead of a single engine. Let independent tool calls
use independent connections so a slow call doesn't block unrelated ones.
Larger change; interacts with the ephemeral-primary / per-session workspace
model and the catalog-bootstrap-once logic. - Cheap read-only fast path. Let
status(and other non-engine-mutating
introspection) answer without taking the engine lock — e.g. from cached
metadata + the daemon health port — so diagnostics never hang even if the
data plane is stalled. - Run blocking calls on
spawn_blockingwith a watchdog that drops/replaces
the engine if an op exceeds a liveness deadline (reusing the existing
ConnectionLost→ drop-and-reconnect path inwith_engine).
Acceptance criteria (rough)
- A stalled or very slow
hyperdoperation on one connection does not make
unrelated tool calls (especiallystatus) hang indefinitely. - Legitimate long-running analytics queries are not aborted by a
duration-based cutoff. - The fix is verified with a test that simulates a wedged/slow connection and
asserts a second concurrent call still returns (or fails fast).
Related
- TCP keepalive fix (the immediate-trigger mitigation): branch
fix/daemon-cold-start-dedup, PR for v0.5.1. - Resident-by-default daemon that increased exposure: #114 / v0.5.0.
- 主要言語
- Rust
- スター
- 3
- フォーク
- 2
- 平均マージ
- 22時間 10分
- マージ済み PR(30日)
- 58
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
tableau/hyper-api-rust のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
tableau/hyper-api-rust#294 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
tableau/hyper-api-rust#311 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
tableau/hyper-api-rust#305 ·
メンテナーはふだん 1 日以内に返信
-
Windows Named Pipe: verify DACL denies other users, and measure read-path perf for MCP workloadsオープン
難易度 4/5 3〜5日 初心者へのやさしさ 38/100
tableau/hyper-api-rust#302 ·
メンテナーはふだん 1 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 72/100
tableau/hyper-api-rust#300 ·
メンテナーはふだん 1 日以内に返信
tableau/hyper-api-rust の issue をすべて見る
似ている issue
-
bug CLI exec tool-calls
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
maintainer-needed p2 triaged ui windows
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
メンテナーはふだん 1 日以内に返信
-
ai_p2
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
ClickHouse/ClickHouse#123351 ·
メンテナーはふだん 1 日以内に返信
-
documentation
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
github/copilot-sdk#2804 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信