Operator throttles itself at scale with >400 CHI pods — reconcile concurrency and the K8s client rate limiter aren't linked
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- go, kubernetes
- Domain
- devops, infrastructure
Research direction
Start by tracing OPERATOR_K8S_CLIENT_QPS_LIMIT, OPERATOR_K8S_CLIENT_BURST_LIMIT, and the reconcile thread and shard-concurrency settings into the Kubernetes client configuration. Reproduce the client-side throttling behavior on a large workload, then define and implement one approach for linking or surfacing the conflicting limits; done means the conflict is visible or prevented and the relevant behavior is verified.
Written by the indexing model from the issue text.
Description
We ran into this on the large CHI in production. It took a while to find because the operator doesn't surface it anywhere.
The operator has two settings that don't know about each other: how much it reconciles in parallel, and how many requests per second it's allowed to send to the Kubernetes API. Turning up concurrency doesn't help if the request rate is low, because all the parallel work still waits on the same rate limit.
The request rate is controlled by OPERATOR_K8S_CLIENT_QPS_LIMIT and OPERATOR_K8S_CLIENT_BURST_LIMIT. If those aren't set, the operator falls back to the Kubernetes client defaults of 5 requests per second with a burst of 10. That's 5 requests per second for the whole operator regardless of how many hosts it manages.
Concurrency is configured separately, through the reconcile thread and shard-concurrency settings. Nothing compares the two. You can tell the operator to reconcile several hundreds shards in parallel while it's still capped at 5 requests per second, and it won't warn you that the two are in conflict.
The impact shows up at scale. A few hundred hosts at roughly ten API calls each is thousands of calls per reconcile, and at 5 requests per second that's about ten minutes of waiting on the rate limit before any useful work happens. With host >400 there was almost no progress in deployment. The only evidence is an info-level log line from the Kubernetes client, which most people aren't watching:
Waited for 9.99s due to client-side throttling, not priority and fairness, request: GET:https://.../api/v1/namespaces/<ns>/pods/<pod>
Confirmed on 0.27.2 (latest upstream) and on our internal 0.25.x builds.
Fix ideas
The goal is to stop the two settings from drifting apart without anyone noticing.
- Check it at startup: if the configured concurrency implies a request rate above the limit, log a warning (or refuse to start) that names both values and the settings to raise.
- Make the throttling visible: raise it to a log level operators see, or expose it as a metric.
- Set defaults that scale with host count instead of inheriting the client's 5/10.
Workaround today
Raise OPERATOR_K8S_CLIENT_QPS_LIMIT and OPERATOR_K8S_CLIENT_BURST_LIMIT on the operator. We used 200/400, which unblocked the internal cluster.
- Dominant language
- Go
- Stars
- 2.6k
- Forks
- 577
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 4
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Altinity/clickhouse-operator
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Altinity/clickhouse-operator#2093 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
Altinity/clickhouse-operator#2094 ·
Maintainers usually reply within 2 days
-
[Regression / Discussion] Loss of hot-reloaded password rotation after removal of k8s_secret_* in 0.27.4Possibly taken @sunsingerus claimed this 3 days ago. Openplanned for review
Difficulty 5/5 Over a week Newbie friendliness 38/100
Altinity/clickhouse-operator#2092 · 1 comment · 1 assignee ·
Maintainers usually reply within 2 days
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
Altinity/clickhouse-operator#2089 ·
Maintainers usually reply within 2 days
-
Difficulty 5/5 Over a week Newbie friendliness 45/100
Altinity/clickhouse-operator#2064 · 3 comments ·
Maintainers usually reply within 2 days
All issues in Altinity/clickhouse-operator
Similar issues
-
[correctness][missing-coverage][sort] Strict uniqueness checks lack numeric-key equivalence coverageOpen
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 69/100
SpecterOps/Janus#6 · 1 comment ·
-
Service process inherits the caller's cwd at first use, holding that folder open on Windows (EBUSY)Open
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
[docs] Media elements cannot load from a custom protocol (video/audio report MEDIA_ERR_SRC_NOT_SUPPORTED)Possibly taken @vst93 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day