Schema Skew Issue
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Tranquilla
- Stack tecnologico
- go, kubernetes
- Ambito
- databases, distributed-systems, infrastructure
Direzione di ricerca
Inizia tracciando migrateTables(), HostWithTablesCreated() e il percorso di reconcile che aggiunge gli host a remote_servers.xml. Esamina come gli errori di migrazione influiscono su CHI.status, sulla disponibilità degli shard e sul comportamento di retry o re-queue, usando come casi gli scenari di bootstrap e /ping riportati. Done dovrebbe definire e verificare un comportamento coerente quando la creazione dello schema o del database replicato fallisce, ma l’issue non specifica quale policy sia prevista.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi team.
We are using the operator . and sometimes due to some race condition we endup with permanent schema skew.
examples: if there was some error while migrateTables() run,the error is simply logged and proceed (Host is pushed and added to HostWithTablesCreated() meaning it will not attempt to fix the schema skew. also the CHI.status transition to completed, and all shards pods mark as ready regardless of the status since they rely on /ping.
Or even if there are 'Create Table .. ON CLUSTER' running on shard-0 while shard-1 is still bootstrap it might miss some of queries since kubelet syncFrequency (by default 1m) will take time till shard-0 aware shard-1 added to the actual remote_servers.xml on Disk.
I wondering what is the best way to solve it since it seems the Operator is 'best-effort' for DDL alignment. I thought to migrate and use Replicated Database engine which solve exactly that. but even if I will switch the Engine, in case of error It seems the Operator is still best-effort for Run the Create Database replicated on new pods/shards.
for example: shard-2 join the cluster the operator will attempt to run 'Create Database Engine = Replicated'. but if also that fail it will add that Host to cluster.
I thought about, maybe not swallow the error on migrateTables(), and raise the error. and maybe re-queue so the reconcile will be retried? or there is reason having the operator working on best-effort strategy?
even if migrateTables succeed or fail we will end-up with Host added to HostWithTabelsCreated(). and also Host added to remote_servers.xml which might get us into permanent schema skew
maybe readiness gate but base on having lag related (DDL/Replication Lag) https://github.com/Altinity/clickhouse-operator/issues/1662 readiness probe will do more harm than good.
We are using versions 0.25.2 and 0.26.3
Thanks
- Lingua principale
- Go
- Stelle
- 2.6k
- Fork
- 577
- Merge medio
- 1g 6h
- PR unite (30g)
- 4
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Altinity/clickhouse-operator
-
Counter metrics are missing the conventional _total suffixForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
Altinity/clickhouse-operator#2093 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
Altinity/clickhouse-operator#2098 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
Altinity/clickhouse-operator#2096 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 72/100
Altinity/clickhouse-operator#2094 ·
I maintainer di solito rispondono entro 2 giorni
-
[Regression / Discussion] Loss of hot-reloaded password rotation after removal of k8s_secret_* in 0.27.4Forse già presa @sunsingerus l’ha presa 4 giorni fa. Apertaplanned for review
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
Altinity/clickhouse-operator#2092 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di Altinity/clickhouse-operator
Issue simili
-
Discriminator mapping keys are listed in a random orderForse già presa @reuvenharrison l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Idle compaction monitors LIST the replica every tick when the newest destination file spans more than one TXIDForse già presa @pishuv l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
benbjohnson/litestream#1563 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
agent-research agent-review-finding chore
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
jordansmall/spindrift#4922 ·
I maintainer di solito rispondono entro 1 giorno
-
gcsartifact: deleting a missing version returns an errorForse già presa @ktsoator l’ha presa oggi. Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 2 giorni