SYSTEM DROP REPLICA being ran from an empty replica if `0-0` is being recovered from its replica
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 58/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- go, kubernetes
- Domain
- databases, infrastructure
Research direction
Start in pkg/controller/chi/worker-deleter.go and inspect how the host for SYSTEM DROP REPLICA is selected during replica cleanup. Reproduce recovery of 0-0 and 0-1 with ClickHouse Keeper, then verify the cleanup works when the first replica lacks Keeper state and add coverage for the recovery scenario.
Written by the indexing model from the issue text.
Description
Hello, while testing disaster recovery scenarios using this operator I noticed that if I have a ClickHouse cluster with 0-0 0-1 1-0 1-1 pods and I delete the PVC and STS of , for example, 0-1 replica and then trigger the operator to reconcile the cluster then the new STS with new PVCs gets recreated and all the tables get sucessfully replicated (I tested using Atomic and Replicated database types)
But the odd behavior I noticed was that if I do the same with 0-0 then the Atomic database and it's tables get recovered from the healthy replica but the Replicated database and tables fail due to broken ClickHouse Keeper state (it's deployed using ClickHouseKeeperInstallation CRD).
I did some investigation into it with the help of an LLM and it pointed out that this function in pkg/controller/chi/worker-deleter.go always drops the stale Zookeeper paths by running SYSTEM DROP REPLICA commands through the *-0 replica which, upon initialization doesn't have that Keeper data.
var hostToRunOn *api.Host
if shard := hostToDrop.GetShard(); shard != nil {
hostToRunOn = shard.FirstHost()
}
Thus it requires more manual input than if I needed to restore the *-1 replicas. So I'm wondering is this an intended behaviour and I'm missing something or is it a bug?
The same agent also suggested a fix but I'm not familiar enough with the codebase to say that it will definitely fix it
var hostToRunOn *api.Host
if shard := hostToDrop.GetShard(); shard != nil {
shard.WalkHosts(func(host *api.Host) error {
if hostToRunOn == nil {
hostToRunOn = host
}
if hostToRunOn == hostToDrop && host != hostToDrop {
hostToRunOn = host
}
return nil
})
}
- Dominant language
- Go
- Stars
- 2.6k
- Forks
- 577
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 4
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Altinity/clickhouse-operator
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Altinity/clickhouse-operator#2093 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
Altinity/clickhouse-operator#2094 ·
Maintainers usually reply within 2 days
-
[Regression / Discussion] Loss of hot-reloaded password rotation after removal of k8s_secret_* in 0.27.4Possibly taken @sunsingerus claimed this 3 days ago. Openplanned for review
Difficulty 5/5 Over a week Newbie friendliness 38/100
Altinity/clickhouse-operator#2092 · 1 comment · 1 assignee ·
Maintainers usually reply within 2 days
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
Altinity/clickhouse-operator#2089 ·
Maintainers usually reply within 2 days
-
Difficulty 5/5 Over a week Newbie friendliness 45/100
Altinity/clickhouse-operator#2064 · 3 comments ·
Maintainers usually reply within 2 days
All issues in Altinity/clickhouse-operator
Similar issues
-
kind/bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 4 days
-
bug needs-acceptance
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
vllm-project/semantic-router#4744 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
jaegertracing/jaeger#9794 ·
Maintainers usually reply within 1 day
-
ScalingModifiers formula fails with "formula returned non-float result" when expression evaluates to an integerPossibly taken @Sarthak-Pandey claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
kedacore/keda#8270 · 1 comment ·
Maintainers usually reply within 1 day