Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Operator throttles itself at scale with >400 CHI pods — reconcile concurrency and the K8s client rate limiter aren't linked

オープン
#2,045 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
48/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
静か
技術スタック
go, kubernetes

調査の方向性

まず、OPERATOR_K8S_CLIENT_QPS_LIMIT、OPERATOR_K8S_CLIENT_BURST_LIMIT、および reconcile スレッドと shard-concurrency の設定が Kubernetes クライアント設定に反映される経路を追跡します。大規模なワークロードでクライアント側のスロットリング動作を再現し、その後、競合する制限を関連付けるか可視化するための1つのアプローチを定義して実装します。競合が可視化または防止され、関連する動作が検証されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

big deployment planned for review research required

We ran into this on the large CHI in production. It took a while to find because the operator doesn't surface it anywhere.

The operator has two settings that don't know about each other: how much it reconciles in parallel, and how many requests per second it's allowed to send to the Kubernetes API. Turning up concurrency doesn't help if the request rate is low, because all the parallel work still waits on the same rate limit.

The request rate is controlled by OPERATOR_K8S_CLIENT_QPS_LIMIT and OPERATOR_K8S_CLIENT_BURST_LIMIT. If those aren't set, the operator falls back to the Kubernetes client defaults of 5 requests per second with a burst of 10. That's 5 requests per second for the whole operator regardless of how many hosts it manages.

Concurrency is configured separately, through the reconcile thread and shard-concurrency settings. Nothing compares the two. You can tell the operator to reconcile several hundreds shards in parallel while it's still capped at 5 requests per second, and it won't warn you that the two are in conflict.

The impact shows up at scale. A few hundred hosts at roughly ten API calls each is thousands of calls per reconcile, and at 5 requests per second that's about ten minutes of waiting on the rate limit before any useful work happens. With host >400 there was almost no progress in deployment. The only evidence is an info-level log line from the Kubernetes client, which most people aren't watching:

Waited for 9.99s due to client-side throttling, not priority and fairness, request: GET:https://.../api/v1/namespaces/<ns>/pods/<pod>

Confirmed on 0.27.2 (latest upstream) and on our internal 0.25.x builds.

Fix ideas

The goal is to stop the two settings from drifting apart without anyone noticing.

  • Check it at startup: if the configured concurrency implies a request rate above the limit, log a warning (or refuse to start) that names both values and the settings to raise.
  • Make the throttling visible: raise it to a log level operators see, or expose it as a metric.
  • Set defaults that scale with host count instead of inheriting the client's 5/10.

Workaround today

Raise OPERATOR_K8S_CLIENT_QPS_LIMIT and OPERATOR_K8S_CLIENT_BURST_LIMIT on the operator. We used 200/400, which unblocked the internal cluster.

主要言語
Go
スター
2.6k
フォーク
577
平均マージ
1日 6時間
マージ済み PR(30日)
4

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Altinity/clickhouse-operator のほかの issue

Altinity/clickhouse-operator の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。