Close connections to a TiDB that the health check has marked down, instead of waiting for TCP to time out
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- go
- Ambito
- backend, networking
Direzione di ricerca
Start by reading TiProxy's health-check handling and backend/client connection lifecycle, especially where a backend is marked down and where processLock governs socket handling. The issue does not name specific files or tests, so locate those entry points and existing failure tests first. Done means an opt-in behavior closes idle connections when a backend goes down and applies the requested failover-timeout behavior to in-flight connections.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Feature Request
Describe your feature request related problem
When a TiDB instance fails without closing its sockets (node/instance failure, network partition, or a pod killed before TiDB can drain), TiProxy detects it (health check marks the backend down) but does nothing to the client connections that are already on that backend.
Measured on TiProxy v1.3.2 + TiDB v8.5.8 on Kubernetes (2 TiDB, 2 TiProxy), client with pool_pre_ping. Test: the EC2 instance hosting tidb-0 was terminated.
| Time (from failure) | Event |
|---|---|
| 0 s | Node hosting tidb-0 terminated; TiDB logs nothing further (no shutdown, no FIN to TiProxy) |
| +14 s | `update backend ... cur="down, err: connect status port failed: ... context deadline exceeded"` |
| +16.3–16.5 s | Short statements sent right after the failure: `backend disconnects ... execute_time=15.45–15.51s cmd=Query`, client gets EOF |
| +46 s | Long query started 17 s before the failure: `backend disconnects ... execute_time=1m3.2s cmd=Query query="select sleep(?)"`, client gets EOF |
Describe the feature you'd like
When the health check marks a TiDB down, we'd like TiProxy to:
- Close idle client connections on that TiDB right away, so the application's next ping fails instantly and its pool reconnects to a healthy TiDB.
- After failover-timeout, also close connections that still have a command in flight. This should break the wait on the TiDB side too (close the backend socket or put a deadline on it), so the client gets an error in seconds rather than a minute.
- Apply the unhealthy keepalive to the backend socket immediately when the TiDB is marked down, without waiting for processLock, so even connections with a stuck command get the shorter timeouts.
Teachability, Documentation, Adoption, Migration Strategy
Keep it off by default so current behaviour doesn't change. It's most useful for OLTP applications behind connection pools, where each pooled connection to a dead TiDB costs one full TCP timeout.
- Lingua principale
- Go
- Stelle
- 73
- Fork
- 41
- Merge medio
- 10h 52m
- PR unite (30g)
- 20
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di pingcap/tiproxy
-
type/enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
-
severity/minor type/bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
I maintainer di solito rispondono entro 1 giorno
-
type/bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 64/100
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di pingcap/tiproxy
Issue simili
-
bug triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
FairwindsOps/nova#484 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
automated-analysis code-quality cookie
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
github/gh-aw#67517 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
[otelcol] print-config help text still requires the removed otelcol.printInitialConfig feature gateForse già presa @girishkvs l’ha presa oggi. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
open-telemetry/opentelemetry-collector#16143 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug good first issue load-balancing
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
ktrubilo9/edge-proxy#53 ·