Caddy crashes under load, with huge number of mutex locks in cache handler
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- go, redis
- Área
- backend, performance
Línea de trabajo
Comienza con la integración de cache-handler y las ubicaciones de los stack traces en github.com/darkweak/souin/pkg/surrogate/providers/common.go y pkg/middleware/middleware.go; después, compáralas con net/http/header.go y caddyhttp/responsewriter.go. Reproduce el escenario de carga usando el Caddyfile y la configuración de caché proporcionados. Se considera terminado cuando el fallo de la map concurrente y la gran cantidad de goroutines esperando a mutexes ya no se produzcan bajo carga.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Under reasonably small load (~250 req/sec), when using the cache-handler (we're using the go-redis storage, but I saw this happen with the default one as well), we get a fatal error: concurrent map iteration and map write in Caddy, this happens reasonably frequently (maybe once a day), requiring a restart.
fatal error: concurrent map iteration and map write
goroutine 27695297 [running]:
net/http.Header.Clone(...)
net/http/header.go:101
net/http.(*response).WriteHeader(0xc920c73c00, 0x1f4)
net/http/server.go:1231 +0x1e7
github.com/caddyserver/caddy/v2/modules/caddyhttp.(*responseRecorder).WriteHeader(0x1ef02e0?, 0x1a70889?)
github.com/caddyserver/caddy/[email protected]/modules/caddyhttp/responsewriter.go:167 +0xb6
github.com/caddyserver/caddy/v2/modules/caddyhttp.(*Server).ServeHTTP(0xc00089a008, {0x1f00d70, 0xc920c73c00}, 0xcf30e57040)
github.com/caddyserver/caddy/[email protected]/modules/caddyhttp/server.go:415 +0x1535
net/http.serverHandler.ServeHTTP({0xcfd3b14750?}, {0x1f00d70?, 0xc920c73c00?}, 0x6?)
net/http/server.go:3210 +0x8e
net/http.(*conn).serve(0xcef56bf4d0, {0x1f04538, 0x1070ae8cba0})
net/http/server.go:2092 +0x5d0
created by net/http.(*Server).Serve in goroutine 181
net/http/server.go:3360 +0x485
This is accompanied by about 1.9 million of these goroutine "Sync.lock.mutex" states, implying (I guess) that a huge number of goroutines are waiting on a mutex.
goroutine 21446179 [sync.Mutex.Lock, 92 minutes]:
sync.runtime_SemacquireMutex(0xc918f18cb0?, 0xb8?, 0xc001a52708?)
runtime/sema.go:95 +0x25
sync.(*Mutex).lockSlow(0xc0008fc640)
sync/mutex.go:173 +0x15d
sync.(*Mutex).Lock(...)
sync/mutex.go:92
github.com/darkweak/souin/pkg/surrogate/providers.(*baseStorage).storeTag(0xc000a80360, {0x0, 0x0}, {0xc61aaf6ac0, 0x38}, 0xcb96172320)
github.com/darkweak/[email protected]/pkg/surrogate/providers/common.go:168 +0xad
github.com/darkweak/souin/pkg/surrogate/providers.(*baseStorage).Store(0xc000a80360, 0xc918f18ec0?, {0xc428c81400?, 0x90?}, {0x19dce40?, 0xcc0cfff7
github.com/darkweak/[email protected]/pkg/surrogate/providers/common.go:232 +0x27f
github.com/darkweak/souin/pkg/middleware.(*SouinBaseHandler).Store.func2({{0x0, 0x0}, 0x12d, {0x0, 0x0}, 0x0, 0x0, 0xca7ca1ea50, {0x1efafc8, 0xcd3d
github.com/darkweak/[email protected]/pkg/middleware/middleware.go:376 +0xcd
created by github.com/darkweak/souin/pkg/middleware.(*SouinBaseHandler).Store in goroutine 21446047
github.com/darkweak/[email protected]/pkg/middleware/middleware.go:375 +0x2891
2024-12-06T08:04:16.889267314Z status_proxy_caddy_caddy.0.t7ydd1vd5hya@statuspage-1 |
My theory is that the crash happens when some pool is finally exhausted.
Attached a log file containing the first 1000 lines, the full thing is many gigabytes (but I can provide if interested, they're all different goroutines waiting at the same point in the code)
Build:
FROM caddy:2.8.4-builder AS builder
RUN xcaddy build \
--with github.com/pberkel/[email protected] \
--with github.com/caddyserver/[email protected] \
--with github.com/darkweak/storages/go-redis/caddy \
--with github.com/aksdb/[email protected]
Caddyfile, some values redacted.
(ip_block) {
@block {
remote_ip REDACTED
}
respond @block "Access denied" 403
}
(host_block) {
@host_block {
host REDACTED
}
respond @host_block "Rate limited" 429
}
{
on_demand_tls {
ask https://REDACTED
}
storage redis {
host "{$REDIS_HOST}"
username "{$REDIS_USER}"
password "{$REDIS_PASSWORD}"
}
servers {
metrics
}
cache {
# how our cache shows in the response header; default is Souin.
cache_name sc
redis {
configuration {
Addrs "{$REDIS_CACHE_HOST}:6379"
User "{$REDIS_CACHE_USER}"
Password "{$REDIS_CACHE_PASSWORD}"
DB 0
}
}
}
order cgi before respond
}
# Catch-all for any domain
:443 {
log
tls {
# we do some funky certificate selection stuff with a local script, fallback to the default
get_certificate http http://localhost:4434/cert
on_demand
}
import ip_block
import host_block
cache {
key {
hide
headers Accept-Language
}
# ignore any cache-control indicators from the client
mode bypass_request
}
# Reverse proxy to server.betteruptime.com
reverse_proxy https://REDACTED {
header_up Host {upstream_hostport}
header_up X-Real-IP {http.request.remote.host}
}
}
http://:9190 {
metrics /metrics
}
http://localhost:4434 {
cache {
stale 5m
ttl 5m
}
cgi /cert /etc/caddy/get_fallback_cert.sh
}
- Lenguaje dominante
- Go
- Estrellas
- 396
- Forks
- 29
- Merge medio
- 32 min
- PR fusionados (30 d)
- 1
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de caddyserver/cache-handler
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
caddyserver/cache-handler#138 · 1 comentario · 3 reacciones ·
-
Admin API purge/invalidation still panics with nil pointer on v0.16.0 (same root cause as #140)Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 76/100
caddyserver/cache-handler#143 · 2 comentarios ·
-
Can not flush cacheAbierto
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
caddyserver/cache-handler#140 · 4 comentarios · 1 reacción ·
-
`Incomming` has a typoAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 45/100
caddyserver/cache-handler#139 · 1 comentario ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
caddyserver/cache-handler#137 · 1 comentario ·
Todos los issues de caddyserver/cache-handler
Issues similares
-
proxy logs "no user in context" at error level for every data gateway downloadPosiblemente ocupada @paul43210 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 83/100
txn2/mcp-data-platform#2063 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 83/100
kubernetes-sigs/kueue#16990 ·
Los mantenedores suelen responder en 1 día
-
enhancement exporter/awss3 needs triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
open-telemetry/opentelemetry-collector-contrib#51905 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
stellar/stellar-horizon#245 ·
Los mantenedores suelen responder en 1 día