Caddy crashes under load, with huge number of mutex locks in cache handler
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 25/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- go, redis
- 领域
- backend, performance
调研方向
先从 cache-handler 集成以及 github.com/darkweak/souin/pkg/surrogate/providers/common.go 和 pkg/middleware/middleware.go 中的堆栈跟踪位置开始,然后将它们与 net/http/header.go 和 caddyhttp/responsewriter.go 进行比较。使用提供的 Caddyfile 和缓存配置重现负载场景。在负载下不再发生并发 map 崩溃以及大量 goroutine 等待 mutex 的情况,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Under reasonably small load (~250 req/sec), when using the cache-handler (we're using the go-redis storage, but I saw this happen with the default one as well), we get a fatal error: concurrent map iteration and map write in Caddy, this happens reasonably frequently (maybe once a day), requiring a restart.
fatal error: concurrent map iteration and map write
goroutine 27695297 [running]:
net/http.Header.Clone(...)
net/http/header.go:101
net/http.(*response).WriteHeader(0xc920c73c00, 0x1f4)
net/http/server.go:1231 +0x1e7
github.com/caddyserver/caddy/v2/modules/caddyhttp.(*responseRecorder).WriteHeader(0x1ef02e0?, 0x1a70889?)
github.com/caddyserver/caddy/[email protected]/modules/caddyhttp/responsewriter.go:167 +0xb6
github.com/caddyserver/caddy/v2/modules/caddyhttp.(*Server).ServeHTTP(0xc00089a008, {0x1f00d70, 0xc920c73c00}, 0xcf30e57040)
github.com/caddyserver/caddy/[email protected]/modules/caddyhttp/server.go:415 +0x1535
net/http.serverHandler.ServeHTTP({0xcfd3b14750?}, {0x1f00d70?, 0xc920c73c00?}, 0x6?)
net/http/server.go:3210 +0x8e
net/http.(*conn).serve(0xcef56bf4d0, {0x1f04538, 0x1070ae8cba0})
net/http/server.go:2092 +0x5d0
created by net/http.(*Server).Serve in goroutine 181
net/http/server.go:3360 +0x485
This is accompanied by about 1.9 million of these goroutine "Sync.lock.mutex" states, implying (I guess) that a huge number of goroutines are waiting on a mutex.
goroutine 21446179 [sync.Mutex.Lock, 92 minutes]:
sync.runtime_SemacquireMutex(0xc918f18cb0?, 0xb8?, 0xc001a52708?)
runtime/sema.go:95 +0x25
sync.(*Mutex).lockSlow(0xc0008fc640)
sync/mutex.go:173 +0x15d
sync.(*Mutex).Lock(...)
sync/mutex.go:92
github.com/darkweak/souin/pkg/surrogate/providers.(*baseStorage).storeTag(0xc000a80360, {0x0, 0x0}, {0xc61aaf6ac0, 0x38}, 0xcb96172320)
github.com/darkweak/[email protected]/pkg/surrogate/providers/common.go:168 +0xad
github.com/darkweak/souin/pkg/surrogate/providers.(*baseStorage).Store(0xc000a80360, 0xc918f18ec0?, {0xc428c81400?, 0x90?}, {0x19dce40?, 0xcc0cfff7
github.com/darkweak/[email protected]/pkg/surrogate/providers/common.go:232 +0x27f
github.com/darkweak/souin/pkg/middleware.(*SouinBaseHandler).Store.func2({{0x0, 0x0}, 0x12d, {0x0, 0x0}, 0x0, 0x0, 0xca7ca1ea50, {0x1efafc8, 0xcd3d
github.com/darkweak/[email protected]/pkg/middleware/middleware.go:376 +0xcd
created by github.com/darkweak/souin/pkg/middleware.(*SouinBaseHandler).Store in goroutine 21446047
github.com/darkweak/[email protected]/pkg/middleware/middleware.go:375 +0x2891
2024-12-06T08:04:16.889267314Z status_proxy_caddy_caddy.0.t7ydd1vd5hya@statuspage-1 |
My theory is that the crash happens when some pool is finally exhausted.
Attached a log file containing the first 1000 lines, the full thing is many gigabytes (but I can provide if interested, they're all different goroutines waiting at the same point in the code)
Build:
FROM caddy:2.8.4-builder AS builder
RUN xcaddy build \
--with github.com/pberkel/[email protected] \
--with github.com/caddyserver/[email protected] \
--with github.com/darkweak/storages/go-redis/caddy \
--with github.com/aksdb/[email protected]
Caddyfile, some values redacted.
(ip_block) {
@block {
remote_ip REDACTED
}
respond @block "Access denied" 403
}
(host_block) {
@host_block {
host REDACTED
}
respond @host_block "Rate limited" 429
}
{
on_demand_tls {
ask https://REDACTED
}
storage redis {
host "{$REDIS_HOST}"
username "{$REDIS_USER}"
password "{$REDIS_PASSWORD}"
}
servers {
metrics
}
cache {
# how our cache shows in the response header; default is Souin.
cache_name sc
redis {
configuration {
Addrs "{$REDIS_CACHE_HOST}:6379"
User "{$REDIS_CACHE_USER}"
Password "{$REDIS_CACHE_PASSWORD}"
DB 0
}
}
}
order cgi before respond
}
# Catch-all for any domain
:443 {
log
tls {
# we do some funky certificate selection stuff with a local script, fallback to the default
get_certificate http http://localhost:4434/cert
on_demand
}
import ip_block
import host_block
cache {
key {
hide
headers Accept-Language
}
# ignore any cache-control indicators from the client
mode bypass_request
}
# Reverse proxy to server.betteruptime.com
reverse_proxy https://REDACTED {
header_up Host {upstream_hostport}
header_up X-Real-IP {http.request.remote.host}
}
}
http://:9190 {
metrics /metrics
}
http://localhost:4434 {
cache {
stale 5m
ttl 5m
}
cgi /cert /etc/caddy/get_fallback_cert.sh
}
- 主要语言
- Go
- 星标
- 396
- 派生
- 29
- 平均合并
- 32 分钟
- 30 天内合并 PR
- 1
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
caddyserver/cache-handler 的其他 Issue
-
难度 1/5 1 小时以内 新手友好度 72/100
caddyserver/cache-handler#138 · 1 条评论 · 3 个 reaction ·
-
难度 3/5 1-2 天 新手友好度 76/100
caddyserver/cache-handler#143 · 2 条评论 ·
-
难度 3/5 1-2 天 新手友好度 55/100
caddyserver/cache-handler#140 · 4 条评论 · 1 个 reaction ·
-
难度 1/5 1 小时以内 新手友好度 45/100
caddyserver/cache-handler#139 · 1 条评论 ·
-
难度 4/5 3-5 天 新手友好度 35/100
caddyserver/cache-handler#137 · 1 条评论 ·
查看 caddyserver/cache-handler 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 88/100
-
[Chore] Remove dead AutogenV2 feature flag可能已有人在做 @geeknishantkyeus 今天认领。 未关闭bug triage
难度 2/5 1-3 小时 新手友好度 75/100
kyverno/kyverno#17936 · 1 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
enhancement model:sonnet phase-2-optimize runtime:claude security size:S
难度 2/5 1-3 小时 新手友好度 62/100
FootprintAI/Containarium#2416 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 66/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
维护者通常 1 天内回复