perAttemptRecvTimeout in RetryPolicy does nothing
维护者通常 1 天内回复
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 55/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
首先跟踪 RetryPolicy 和 RetriableStream,重点关注 perAttemptRecvTimeoutNanos 如何到达 makeRetryDecision。使用已报告的黑洞场景,或使用等效的 channel 和 RetryPolicy,验证超过配置超时的尝试会产生 DEADLINE_EXCEEDED,并触发预期的重试行为。
由索引模型根据 Issue 内容生成。
描述
What version of gRPC-Java are you using?
1.76.0 (but it's the same problem in the latest release too)
What is your environment?
Android app with gprc-java and grpc-android libs
What did you expect to see?
https://github.com/grpc/grpc-java/pull/8301 added perAttemptRecvTimeoutNanos to the RetryPolicy but the problem that it's never used so it's not possible to set a timeout for an attempt within the gRPC call. Not sure but might be related to this issue https://github.com/grpc/grpc-java/issues/1943
What did you see instead?
I would expect RetriableStream to use perAttemptRecvTimeoutNanos from RetryPolicy so makeRetryDecision will actually retry when an attempt takes longer than specified in perAttemptRecvTimeoutNanos
Steps to reproduce the bug
Can be reproduced with any channel and RetryPolicy that has perAttemptRecvTimeout.
A bit of context why do I even need it
I'm looking for a way to detect and recover from a black hole gRPC connection that happens in my app.
I have logs from production with DEADLINE_EXCEEDED exception that either has waiting_for_connection or remote_addr=/10.0.2.2:8443 (it's from local test but the point that there is a remote_addr in production). As far as I understand in the first case the channel is in CONNECTING state and in the second is in READY but neither of them can succeed and backend has no errors. I reproduced the issue locally by using toxiproxy and simulated a black hole connection. My idea is to pass perAttemptRecvTimeout in RetryPolicy so I can rely on gRPC internal retry mechanism and when an attempt fails with DEADLINE_EXCEEDED to call channel.enterIdle (same as AndroidChannelBuilder does when detects the change in network to force a new connection for the next call) in ClientStreamFactory (same idea as for refreshing an expired auth token but without CallCredentials as was suggested here and it actually works in our app https://github.com/grpc/grpc-java/issues/7345#issuecomment-679295003)
object : ClientStreamTracer.Factory() {
override fun newClientStreamTracer(
info: ClientStreamTracer.StreamInfo,
headers: Metadata,
): ClientStreamTracer = object : ClientStreamTracer() {
override fun streamClosed(status: Status) {
if (status.code == Status.Code.DEADLINE_EXCEEDED) {
channel.enterIdle()
}
}
}
}
I also considered keepAlive option but it doesn't seems to work as smooth as the idea above and if I'm not wrong - it will not recover from a black hole connection when the channel is in CONNECTING because it's only sent for an established connection so the channel has to be in READY state. And also it will drain battery and spam the backend.
Also I thought to have a retry interceptor but I'm afraid it's a fragile approach that previously caused us a lot of crashes when it was attempted for token refresh and retry afterwards. (Probably it's possible to implement it flawlessly but it likely it will be difficult to understand and easy to break in the future and therefore not considered).
And as a last resort it's also possible to use try/catch and retry on the call-site and manage the channel there but it requires to change every single call-site in the app and can be easily forgotten for new calls. So not a sustainable option.
Just wanted to share what I already thought of and I believe the first suggestion that uses perAttemptRecvTimeout is the best among listed approaches. I would be happy to hear if there are other possible robust ways to fix it.
- 主要语言
- Java
- 星标
- 12.1k
- 派生
- 4k
- 平均合并
- 2 天 6 小时
- 30 天内合并 PR
- 26
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
grpc/grpc-java 的其他 Issue
-
channelz: ServerData `calls_failed` counter not incremented upon client cancellation可能重新可做 @stdcout42 于 12 天前认领,目前没有进行中的 PR。 未关闭enhancement
难度 2/5 1-3 小时 新手友好度 88/100
grpc/grpc-java#13063 · 4 条评论 ·
维护者通常 1 天内回复
-
Android: ProxyDetectorImpl crashes when DefaultProxySelector contains an invalid proxy port可能已有人在做 @kkmurthyt21 于 27 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 78/100
grpc/grpc-java#13052 · 4 条评论 ·
维护者通常 1 天内回复
-
docs enhancement
难度 2/5 1-3 小时 新手友好度 72/100
grpc/grpc-java#10824 · 8 条评论 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 45/100
维护者通常 1 天内回复
-
binder: RuntimeException during outbound serialization leaks the call and never notifies the peer可能已有人在做 @mvanhorn 今天认领。 未关闭bug
难度 4/5 3-5 天 新手友好度 48/100
grpc/grpc-java#13093 · 1 条评论 ·
维护者通常 1 天内回复
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 60/100
维护者通常 1 天内回复
-
[BUG] S3 CORS responses omit Access-Control-Allow-Credentials for matched origins可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
securityHeaders replaces a route's own Content-Security-Policy (0.9.9; weakens embedders' pages)未关闭bug
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
sqlcipher/sqlcipher-android#97 · 1 条评论 ·
-
area-integrations
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复