Operator stops reconciling permanently after its watch connections drop while idle
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- kubernetes, rust
- Domain
- devops, infrastructure
Research direction
Start with the controller watch and backoff behavior in kube-runtime/src/controller/mod.rs around the referenced line, then inspect how this operator configures Controller and watcher::Config. Reproduce the idle watch disconnect and compare the constant-backoff and reduced-timeout approaches described in the issue. Done means affected operators reconnect after dropped idle watches and continue reconciling without a pod restart.
Written by the indexing model from the issue text.
Description
Affected Stackable version
26.7
Affected Trino version
481
Current and expected behavior
The trino, hive and airflow operators flip from recoverable watch-stream errors to a permanent failed to start watching object: client error (Connect) loop and never reconcile again until the pod is restarted. While wedged the operator looks healthy to Kubernetes: the process is running, its conversion webhook still answers.
Expected
A controller whose watch connection is dropped re-establishes it and continues reconciling, as it does for the first several drops.
Actual
ERROR kube_client::client::builder: failed with error client error (Connect)
WARN stackable_operator::logging::controller: Queued reconcile resulted in an error
controller.name="trinocluster.trino.stackable.tech"
error=failed to start watching object: ServiceError: client error (Connect)
error.sources=[ServiceError: client error (Connect), client error (Connect), deadline has elapsed]
The problems seems to be upstream in kube-rs: https://github.com/kube-rs/kube/blob/main/kube-runtime/src/controller/mod.rs#L1704
Possible explanation for this issue: All watch triggers get merged into a single stream that then gets joint backoff - so if an operator watches e.g. 7 resources, the backoff fires instantly 6 times, which quickly leads to high sleep durations that ultimately go beyond the timeout.
While kube-rs has a long timeout configured (>200 secs) on the observed cluster some kind of proxy for the apiserver seems to terminate idle watches after 60 seconds, which is reliably triggered for the mentioned operators.
Possibly related: https://github.com/kube-rs/kube/issues/1915
As mentioned in the beginning, this potentially affects all our operators
Possible solution
Possible Workarounds:
Smaller, constant backoff
use kube::runtime::utils::Backoff;
struct ConstantBackoff(Duration);
impl Iterator for ConstantBackoff {
type Item = Duration;
fn next(&mut self) -> Option<Duration> { Some(self.0) }
}
impl Backoff for ConstantBackoff {
fn reset(&mut self) {}
}
Controller::new(api, watcher::Config::default())
.trigger_backoff(ConstantBackoff(Duration::from_secs(2)))
Reduce timeout on watcher side:
Maybe as an optional parameter or something.
const WATCH_TIMEOUT_SECS: u32 = 40;
fn watch_config() -> watcher::Config {
watcher::Config::default().timeout(WATCH_TIMEOUT_SECS)
}
Additional context
No response
Environment
No response
Would you like to work on fixing this bug?
None
- Dominant language
- Rust
- Stars
- 63
- Forks
- 14
- Avg merge
- 1d 8h
- Merged PRs (30d)
- 10
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from stackabletech/trino-operator
-
customer-request
Difficulty 2/5 1-3 hours Newbie friendliness 60/100
stackabletech/trino-operator#499 ·
Maintainers usually reply within 1 day
-
bug: pod spec doesn't match sts template specPossibly taken @razvan claimed this 217 days ago. Openrelease-note
stackabletech/trino-operator#854 · 3 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
stackabletech/trino-operator#849 ·
Maintainers usually reply within 1 day
-
customer-request type/feature-improvement
Difficulty 3/5 1-2 days Newbie friendliness 35/100
stackabletech/trino-operator#813 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
stackabletech/trino-operator#806 ·
Maintainers usually reply within 1 day
All issues in stackabletech/trino-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 5 days
-
state:triage-needed
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
ktuner keeps a stale ledger path and can never restore that entryPossibly taken @Frun1na claimed this today. Opencomponent:ktuner
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
agentic-os-org/ANOLISA#6483 · 1 comment ·
Maintainers usually reply within 1 day
-
[Resource]: snapbackOpenresource-submission validation-passed
Difficulty 1/5 Under an hour Newbie friendliness 85/100
hesreallyhim/awesome-claude-code#3095 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day