Centralize retry/backoff configuration on `backon`; inject retries where components give up
维护者通常 4 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 重构
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- rust
调研方向
从 crates/app/src/retry.rs 和 crates/core/src/expbackoff.rs 开始,然后跟踪所列出的 scheduler、relay/dial.rs、quic_upgrade.rs、eth1wrap、dkg、recast 和 wire sink 入口点。在决定应如何分层共享策略之前,比较它们的重试和轮询行为。完成的标准是处理重复的配置和手动循环,并且指定的放弃位置会按照预期的截止时间重试。
由索引模型根据 Issue 内容生成。
描述
Summary
backon is already the workspace retry engine and app/src/retry.rs (the Charon app/retry port) is built on it — but the workspace has drifted into at least 5 distinct backoff parameter sets and 4 drive mechanisms:
retry.rsdefaults (250ms/12s/1.6) — zero call sites (#534 tracks wiring it into the duty callbacks).core::expbackofffast()(100ms/5s) anddefault()(1s/120s) — used by scheduler/sse/bootnode, sometimes driven manually via.next()+expect.p2p/src/relay/dial.rs#L104-L126— a hand-written duplicate ofexpbackoff::default()(its doc even cites the same Charon config), reimplemented because p2p needs a pollableDurationrather than an async wrapper.quic_upgrade.rs— a third scheme in units of minutes (1→512, doubling).eth1wrap— Alloy'sRetryBackoffLayer::new(10, 1000, 100), a fourth policy in a foreign library's units.- Fixed-delay loops in
dkg/sync(250ms),cli test/peers(for attempt in 0..5, 5s intervals). Same family:dkgbusy-polls node signatures on a 100ms ticker, re-locking and cloning the accumulated slot vector every tick (nodesigs.rs#L144-L164), andexchanger.rs#L561polls with a bare 100mssleep— both want aNotify/watchsignal instead of a poll.
Meanwhile, components that should retry just give up and wait for the next tick: scheduler resolve_duties errors are logged "(retrying next slot)" — a beacon-node blip loses a slot's duty resolution; bcast/recast retries next epoch; the wire-layer store/broadcast sinks log and swallow errors (wire.rs#L739-L746 and siblings) — precisely the points Charon wraps in async-retry.
Proposed change
- One retry module (the existing
app::retry+core::expbackoff, merged or clearly layered) exposing the named policies (fast,default, plus a pollable-Durationhelper for poll-based behaviours sorelay/dial.rscan delete its copy). - Replace the manual
.next()/hand-rolled loops withbackon'sRetryableor the shared pollable helper; expresseth1wrap's policy in the same config vocabulary. - Inject retries at the give-up sites above (scheduler duty resolution, recast, wire sinks), with per-duty deadlines from the existing
DeadlineCalculatorplumbing. #534 covers the five Charon duty-callback wrap points; this issue covers the config unification and the remaining sites.
- 主要语言
- Rust
- 星标
- 8
- 派生
- 6
- 平均合并
- 4 天 3 小时
- 30 天内合并 PR
- 18
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NethermindEth/pluto 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 52/100
NethermindEth/pluto#719 ·
维护者通常 4 天内回复
-
enhancement rust
难度 4/5 3-5 天 新手友好度 45/100
NethermindEth/pluto#640 ·
维护者通常 4 天内回复
-
rust
难度 5/5 一周以上 新手友好度 45/100
NethermindEth/pluto#638 · 1 条评论 ·
维护者通常 4 天内回复
-
enhancement rust
难度 5/5 一周以上 新手友好度 35/100
NethermindEth/pluto#635 ·
维护者通常 4 天内回复
-
Improve `alpha test peers`: match Charon's probes; reduce timeouts可能已有人在做 @varex83agent 于 8 天前认领。 未关闭bug track:orchestration-cli
难度 5/5 一周以上 新手友好度 28/100
NethermindEth/pluto#632 ·
维护者通常 4 天内回复
查看 NethermindEth/pluto 的全部 Issue
相似的 Issue
-
✨ enhancement needs-discussion
难度 1/5 1 小时以内 新手友好度 85/100
-
area:docs documentation good first issue priority:low
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
triage:accepted
难度 2/5 1-3 小时 新手友好度 65/100
open-telemetry/otel-arrow#4343 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
mishraprafful/multihull#150 ·
维护者通常 1 天内回复
-
area:tooling bug good first issue priority:P3
难度 2/5 1-3 小时 新手友好度 72/100
michaelnavazhylau/ngspice-rs#129 ·
维护者通常 1 天内回复