Add is_alive liveness-probe API across SDKs (decouple liveness from typed ping deserialization)
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 48/100
- Issue 类型
- 功能
- 描述清晰度
- 描述清楚
- 活跃度
- 冷清
调研方向
从 rust/src/lib.rs:1676、nodejs/src/client.ts:1032、go/client.go:1317、dotnet/src/Client.cs:874 中列出的 Client ping 入口以及对应的 Python client 开始,然后检查 @github/copilot/schemas/api.schema.json。运行 SDK 测试,并为 ping body 无法通过严格反序列化的有效 JSON-RPC 响应添加覆盖。完成的标准是每个 Client 都有适合相应语言的 liveness 方法,现有的 typed ping API 保持不变,并且每个 SDK 都记录这一差异。
由索引模型根据 Issue 内容生成。
描述
Summary
The Copilot CLI host (github/github-app) had to work around a deserialization failure in our SDK's typed ping() API on the session-resume / warm-CLI-pool / liveness-probe paths. See github/github-app#5461 for the workaround.
Their fix introduces a ping_cli_compat(&client) helper that drops to the raw client.call("ping", json!({})) so the result body is never deserialized — they only care whether the JSON-RPC round trip succeeded.
We should give them (and everyone doing liveness checks) a first-class primitive instead of forcing them to bypass our typed API.
Root cause
Across all five SDKs, ping() deserializes the response into a typed body:
- Rust:
Client::ping(&self) -> Result<PingResponse, Error>(rust/src/lib.rs:1676) - Node:
client.ping()returning{message, timestamp, protocolVersion?}(nodejs/src/client.ts:1032) - Go:
client.Ping(ctx, msg) (*PingResponse, error)(go/client.go:1317) - .NET:
client.PingAsync(...) : Task<PingResponse>(dotnet/src/Client.cs:874) - Python: equivalent typed return
The authoritative PingResult schema (@github/copilot/schemas/api.schema.json) requires message, timestamp (date-time string), and protocolVersion (integer > 0). The Rust hand-written PingResponse softens this with #[serde(default)] and Option<u32>, but that only helps for missing fields — wrong types (e.g. null timestamp, integer-shape drift) still fail. Same brittleness exists across the other SDKs.
ping is used by hosts as a liveness check (warm-pool reuse, resumed-session aliveness, retrier "is the cached client still alive" probes). These paths straddle CLI version boundaries — a resumed older CLI process can answer with a slightly different ping body shape, and the entire liveness check fails even though the RPC itself succeeded.
This is exactly the wrong failure mode for a health check: the consumer asked "is the CLI reachable?" and we answered "no" because of a body-shape mismatch.
Proposed fix
Add a dedicated liveness-probe API to every SDK Client:
- Sends the
pingJSON-RPC call. - Returns success based solely on JSON-RPC success — never deserializes the result body.
- Composable with caller-supplied timeouts — no baked-in timeout, so each host (startup probe, warm-pool, resume probe, background keepalive) sets its own budget.
ping()and generatedrpc.ping()stay strict and schema-faithful. Callers who actually want the typed data keep getting it; schema drift continues to surface there as a real error.
Per-language names:
| SDK | API |
|---|---|
| Rust | Client::is_alive(&self) -> bool |
| Node | client.isAlive(): Promise<boolean> |
| Python | client.is_alive() -> bool |
| Go | client.IsAlive(ctx context.Context) bool |
| .NET | client.IsAliveAsync(ct) : Task<bool> |
After this lands, the github/github-app workaround helper goes away and call sites become e.g. existing.is_alive().await.
Rejected alternatives
- Loosen
ping()itself. Considered and rejected.ping()is a typed schema-backed API; if it silently swallows malformed bodies, the contract becomes ambiguous (did the caller want liveness, or the ping data?). It also masks real CLI/schema drift in the API most likely to catch it. - Loosen the generated
rpc.ping(PingRequest). Same reasoning, more so — generated APIs must remain schema-faithful.is_aliveis the explicit escape hatch. - "Fix it only in the CLI." The CLI should still honor the schema, and we should investigate any drift. But liveness checks fundamentally straddle version boundaries, so the SDK needs a body-agnostic primitive regardless.
Acceptance criteria
- New
is_alive/IsAlive/isAlivemethod on the Rust, Node, Python, Go, and .NETClienttypes. - Method calls the
pingJSON-RPC method and returns success purely on RPC success, ignoring the response body. - Existing
ping()/rpc.ping()typed APIs unchanged. - Docs in each SDK clearly distinguish "use
is_alivefor liveness/warm-pool/resume checks" from "useping()when you want the typed response." - Tests cover the case where the CLI returns a
pingbody that fails strict deserialization but is otherwise a valid JSON-RPC success —is_alivereturnstrue,ping()returns an error. - Once shipped, the workaround in github/github-app#5461 is removed.
Design review
Reviewed and agreed with GPT-5.5; consensus on:
- Separate
is_aliveAPI (not looseningping). - Keep generated
rpc.pingstrict. - No baked-in timeout; caller composes timeouts.
- Name
is_aliveclearly communicates intent ("RPC round trip succeeded") and beatsping_raw/ping_check.
- 主要语言
- Java
- 星标
- 10.5k
- 派生
- 1.5k
- 平均合并
- 1 天 9 小时
- 30 天内合并 PR
- 130
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
github/copilot-sdk 的其他 Issue
-
agentic-workflows
难度 2/5 1-3 小时 新手友好度 65/100
github/copilot-sdk#2760 ·
-
难度 2/5 1-3 小时 新手友好度 65/100
github/copilot-sdk#2759 ·
-
documentation
难度 1/5 1 小时以内 新手友好度 85/100
github/copilot-sdk#2758 ·
-
agentic-workflows
难度 2/5 1-3 小时 新手友好度 68/100
github/copilot-sdk#2709 · 1 条评论 ·
-
难度 1/5 1 小时以内 新手友好度 78/100
github/copilot-sdk#2673 ·
查看 github/copilot-sdk 的全部 Issue
相似的 Issue
-
certification
难度 1/5 1 小时以内 新手友好度 80/100
-
难度 2/5 1-3 小时 新手友好度 75/100
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's 未关闭bug ecr
难度 2/5 1-3 小时 新手友好度 75/100
-
Needs: Triage Type: Feature request
难度 2/5 1-3 小时 新手友好度 70/100
AntennaPod/AntennaPod#8794 ·
-
awaiting triage bug Causes friction Hop Gui P1 P2 Transforms
难度 2/5 1-3 小时 新手友好度 75/100