Make leader-for-life leader election more integrated with controller-runtime
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
- 技术栈
- go
调研方向
从 controller-runtime 的 manager.Start() 序列以及 issue 中描述的 leader-for-life 集成开始。跟踪 leader election、liveness 和 readiness probe,以及 controller 启动之间如何交互。完成的标准是 controller-runtime 支持可插拔的 leader-election 实现,使得 leader-for-life 可以在升级期间使用,而不会出现 probe deadlock。
由索引模型根据 Issue 内容生成。
描述
Feature Request
Is your feature request related to a problem? Please describe.
Yes. It isn't possible to use leader-for-life leader election with controller-runtime's manager when also using liveness and readiness probes.
Using controller-runtime's manager out of the box, the following sequence of events happens when manager.Start() is called:
- Liveness and readiness probes are started
- Leader election is started.
- Controllers are started.
When using leader-for-life from this repo, it must be called prior to manager.Start() since controller-runtime doesn't support pluggable leader election implementations. The sequence of events in this case is:
- Leader election is started.
- Liveness and readiness probes are started
- Controllers are started.
Notice that 1) and 2) are swapped. This swap causes deadlocks when upgrading operator deployments that use leader-for-life. When the deployment is attempting to rollout a new version, the new pod starts up and first attempts to become the leader, failing indefinitely until the old pod relinquishes ownership. However the old pod will not relinquish ownership until it disappears and it won't disappear until the new pod reports that it's healthy. Unfortunately the new pod will never be able to report that it's healthy because it needs to be the leader before it starts its liveness and readiness probe servers.
Describe the solution you'd like
To work upstream to make controller-runtime support a pluggable leader election implementation such that leader-for-life can be used by the manager.
- 主要语言
- Go
- 星标
- 40
- 派生
- 42
- 平均合并
- 4 天 46 分钟
- 30 天内合并 PR
- 7
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
operator-framework/operator-lib 的其他 Issue
-
Prune package?未关闭
难度 3/5 1-2 天 新手友好度 25/100
-
难度 5/5 一周以上 新手友好度 25/100
-
lifecycle/frozen
难度 4/5 3-5 天 新手友好度 35/100
operator-framework/operator-lib#103 · 6 条评论 ·
查看 operator-framework/operator-lib 的全部 Issue
相似的 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
open-telemetry/opentelemetry-go-compile-instrumentation#1467 ·
维护者通常 3 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
modelcontextprotocol/go-sdk#1367 · 1 条评论 ·
维护者通常 1 天内回复
-
Python 3.15 support可能已有人在做 @amnesiaof 今天认领。 未关闭L: python L: python:uv
难度 2/5 1-3 小时 新手友好度 72/100
dependabot/dependabot-core#16524 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
duplication
难度 2/5 1-3 小时 新手友好度 78/100
openvibely/openvibely#1443 ·
维护者通常 2 天内回复