Distinguish high load and unavailability of TiDB
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
Línea de trabajo
Empieza leyendo las implementaciones de health-check y router, centrándote en el tiempo de espera de 2 segundos, las tres comprobaciones fallidas y la migración de conexiones todas a la vez descritas en el issue. Reproduce o razona sobre los escenarios de alta carga para dos instancias de TiDB y, después, define y valida un comportamiento que distinga la lentitud temporal de la falta de disponibilidad sin infringir graceful-wait-before-shutdown.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Development Task
When the load of TiDB is very high, it may not respond within 2 seconds, and thus TiProxy treats it as down and migrates ALL connections away from it. The migration is typically very fast and may finish most of them before the next health check.
If there are 2 TiDB instances, A and B. The CPU usage of A is 100% and B is 90%. Theoretically, this may happen:
- TiProxy can not dial A within 2 seconds and treats A as down and B as alive
- TiProxy wants to migrate all connections from A to B
- After 3 seconds, the CPU usage of A becomes 90% and B becomes 100%
- Again, TiProxy thinks A is alive but B is down
- TiProxy wants to migrate all connections from B to A
This situation is just theoretically possible but I'm not sure if it will happen in the real world:
- If the CPU usage of A is 100%, is it possible to migrate many connections within 3 seconds?
- When the CPU usage of B becomes 100%, it's slow to dial B. Will the connection migration be that fast?
Anyway, the strategies of the health check and router are too aggressive:
- The health check treats the backend as down when it just fails in 2 seconds for 3 times. However, I also want to make the health check fast enough so that the
graceful-wait-before-shutdowncan be configured shorter. - The router tries to migrate all connections all at once. Exactly, 10ms for 10 connections and thus 3s for 3000 connections for each TiProxy. 3000 connections almost mean all the connections on one TiDB. However, I also want to make the migration ASAP because it needs to finish within
graceful-wait-before-shutdown.
- Lenguaje dominante
- Go
- Estrellas
- 73
- Forks
- 41
- Merge medio
- 12 h 38 min
- PR fusionados (30 d)
- 21
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de pingcap/tiproxy
-
contribution first-time-contributor type/enhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
pingcap/tiproxy#1239 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
type/enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
Los mantenedores suelen responder en 1 día
-
Flaky test TestForwardUntilErrorAbiertoseverity/minor type/bug
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
-
type/bug
Dificultad 3/5 1-2 días Aptitud para principiantes 64/100
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 38/100
Los mantenedores suelen responder en 1 día
Todos los issues de pingcap/tiproxy
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
golang/go#82033 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 90/100
FootprintAI/Containarium#2338 ·
Los mantenedores suelen responder en 1 día
-
[Bug]: core doesn't build standalone on dev since a17068054 (go-mp3 require dropped, go.sum pruned)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
SagerNet/sing-openvpn#11 ·
-
priority: P3 type: devops
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
jiegui2025/hwspec#57 ·
Los mantenedores suelen responder en 1 día