Distinguish high load and unavailability of TiDB
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
Direzione di ricerca
Inizia leggendo le implementazioni di health-check e router, concentrandoti sul timeout di 2 secondi, sui tre check falliti e sulla migrazione simultanea delle connessioni descritti nell’issue. Riproduci o analizza gli scenari di carico elevato per due istanze TiDB, quindi definisci e convalida un comportamento che distingua la lentezza temporanea dall’indisponibilità senza violare graceful-wait-before-shutdown.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Development Task
When the load of TiDB is very high, it may not respond within 2 seconds, and thus TiProxy treats it as down and migrates ALL connections away from it. The migration is typically very fast and may finish most of them before the next health check.
If there are 2 TiDB instances, A and B. The CPU usage of A is 100% and B is 90%. Theoretically, this may happen:
- TiProxy can not dial A within 2 seconds and treats A as down and B as alive
- TiProxy wants to migrate all connections from A to B
- After 3 seconds, the CPU usage of A becomes 90% and B becomes 100%
- Again, TiProxy thinks A is alive but B is down
- TiProxy wants to migrate all connections from B to A
This situation is just theoretically possible but I'm not sure if it will happen in the real world:
- If the CPU usage of A is 100%, is it possible to migrate many connections within 3 seconds?
- When the CPU usage of B becomes 100%, it's slow to dial B. Will the connection migration be that fast?
Anyway, the strategies of the health check and router are too aggressive:
- The health check treats the backend as down when it just fails in 2 seconds for 3 times. However, I also want to make the health check fast enough so that the
graceful-wait-before-shutdowncan be configured shorter. - The router tries to migrate all connections all at once. Exactly, 10ms for 10 connections and thus 3s for 3000 connections for each TiProxy. 3000 connections almost mean all the connections on one TiDB. However, I also want to make the migration ASAP because it needs to finish within
graceful-wait-before-shutdown.
- Lingua principale
- Go
- Stelle
- 73
- Fork
- 41
- Merge medio
- 12h 38m
- PR unite (30g)
- 21
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di pingcap/tiproxy
-
contribution first-time-contributor type/enhancement
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
pingcap/tiproxy#1239 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
type/enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
-
severity/minor type/bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
I maintainer di solito rispondono entro 1 giorno
-
type/bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 64/100
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di pingcap/tiproxy
Issue simili
-
type/bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
apimachinery yaml: YAMLOrJSONDecoder drops a trailing document shorter than 4 bytesForse già presa @HosniBelfeki l’ha presa oggi. Apertaneeds-triage sig/api-machinery
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
kubernetes/kubernetes#142651 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
kubernetes-sigs/kueue#16649 ·
I maintainer di solito rispondono entro 1 giorno
-
bug(services): classifySBOMStatus misclassifies UnsupportedSchema as generic ReasonSBOMGenerationFailedForse già presa @bhuvan-somisetty l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno