Deprecate Leader for life based leader election
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
- 技術スタック
- go
調査の方向性
まず leader-for-life の実装と、それが manager および controller-runtime とどのように統合されているかを確認します。非推奨化の方針を決定する前に、リンクされている operator-lib と controller-runtime の issue を確認してください。leader-for-life オプションが deprecated になり、今後の Operator SDK リリースでの削除が予定され、upstream 統合に関する問題が解決されていれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Feature Request
Is your feature request related to a problem? Please describe.
Currently, the repository contains a leader-for-life based election model that ensures that a single leader is elected for life during a HA state.
To briefly describe what leader for life and leader for lease approaches are:
(1) Leader for life: A leader pod is selected for life until its garbage collected. This ensures there is only one leader at any instant of time.
(2) Leader for lease: Leader election happens periodically at a defined interval which can be tweaked. When the current leader is not able to renew the lease, a new leader is elected.
More details on both of these approaches are available here.
Approach (2) is implemented upstream, in client-go and is scaffolded by default for the past 30+ releases of SDK, since it is implemented as a part of setting up the manager in controller-runtime. It guarantees faster election of leaders, less downtime, recovery from disconnected/frozen node failures. However, it does not eliminate the split brain scenario - where more than a single leader is available at an instant of time.
Approach (1) on the other hand, was developed long before we had a leader-for-lease implemented upstream. Though it solves the split brain scenario, it does not guarantee recovery from node failure nor faster recovery. We also have issues with integrating it to controller-runtime (https://github.com/operator-framework/operator-lib/issues/48). Neither is it being maintained nor used as widely as leader for lease.
Here is a detailed comment explaining the preference of (2) over the other, wherein users would prefer a faster recovery even though there is split brain scenario intermittently, rather than an implementation that does not guarantee faster recovery.
Describe the solution you'd like
Since leader for life approach is not being widely used, neither works with controller-runtime seamlessly, it is better to adopt a well tested upstream library than to depend on what is currently available as an option.
The solution for this is:
- To deprecate and remove leader for life in future releases of Operator SDK.
- To bring this up upstream (in controller-runtime), for easier integration. This had already been brought up upstream (https://github.com/kubernetes-sigs/controller-runtime/issues/1963) but there was no response on the same.
- 主要言語
- Go
- スター
- 40
- フォーク
- 42
- 平均マージ
- 4日 46分
- マージ済み PR(30日)
- 7
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
operator-framework/operator-lib のほかの issue
-
Prune package?オープン
難易度 3/5 1〜2日 初心者へのやさしさ 25/100
-
lifecycle/frozen
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
operator-framework/operator-lib#103 · コメント 6 件 ·
-
lifecycle/frozen
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
operator-framework/operator-lib#48 · コメント 4 件 ·
operator-framework/operator-lib の issue をすべて見る
似ている issue
-
proxy logs "no user in context" at error level for every data gateway download対応中かも @paul43210 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
txn2/mcp-data-platform#2063 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 83/100
kubernetes-sigs/kueue#16990 ·
メンテナーはふだん 1 日以内に返信
-
enhancement exporter/awss3 needs triage
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
open-telemetry/opentelemetry-collector-contrib#51905 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
stellar/stellar-horizon#245 ·
メンテナーはふだん 1 日以内に返信