0.7.2: fatal error: concurrent map writes in whatsmeowService.StartClient (kills the whole process)
Maintainer thường phản hồi trong vòng 5 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 62/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- go
- Lĩnh vực
- authentication, backend
Hướng nghiên cứu
Start in pkg/whatsmeow/service/whatsmeow.go at the map deletion on line 581, then trace the recursive StartClient retry at line 629 and StartInstance at line 2407. Reproduce concurrent /instance/connect calls while pairing by QR, with handleQRCodes active. Done means the race no longer produces a fatal concurrent map write or process-wide exit, and QR pairing can complete.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Running evoapicloud/evolution-go:0.7.2 with ~12 instances, the process dies repeatedly with an unrecoverable Go runtime error. Not a recoverable panic — it takes down every instance in the deployment. We are at 168 restarts in 54 days.
Stack trace
fatal error: concurrent map writes
goroutine 18577 [running]:
internal/runtime/maps.fatal({0x195b5e4?, 0x0?})
/usr/local/go/src/runtime/panic.go:1046 +0x18
internal/runtime/maps.(*Map).Delete(0xc000ffe9c0, 0x171eb60, 0xc003c3b228)
/usr/local/go/src/internal/runtime/maps/map.go:652 +0x4c
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
/build/pkg/whatsmeow/service/whatsmeow.go:581 +0x28bc
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
created by github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartInstance in goroutine 18558
/build/pkg/whatsmeow/service/whatsmeow.go:2407 +0xa2d
The crashing write is a map delete at whatsmeow.go:581, reached from StartClient. Note StartClient appears recursively via whatsmeow.go:629 — it re-enters itself on retry, so a single instance can have several StartClient frames in flight while other instances are starting from StartInstance in their own goroutines. The map mutated at :581 looks like shared state across instances with no mutex.
When it happens
Every crash we have captured has a live handleQRCodes goroutine in the same dump:
goroutine 72775 [sleep]:
time.Sleep(0xdf8475800)
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.(*MyClient).handleQRCodes.func1()
/build/pkg/whatsmeow/service/whatsmeow.go:785 +0x85
created by github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.(*MyClient).handleQRCodes in goroutine 72773
/build/pkg/whatsmeow/service/whatsmeow.go:704 +0x86
So the trigger for us is: pair an instance by QR while another instance is (re)connecting. The pairing never reaches the auth store — after the restart the instance has an empty JID and needs a new QR. From the outside it looks like "we scanned the QR and it dropped by itself".
Reproduction
- Have several instances registered, some offline.
- Call
POST /instance/connecton more than one of them at roughly the same time (an external reconnect loop is enough). - Pair one instance by QR while that is happening.
It is a race, so it does not reproduce every time — for us it lands about 3 times a day.
Environment
evoapicloud/evolution-go:0.7.2(sha256:6fa601464bd76d3d19ece4c48d8bb5373d025d95d45b939ae4bf0c77b03f5aaa)- whatsmeow
v0.0.0-20260630180629-b572e5bcb92b - Auth store in Postgres (
POSTGRES_AUTH_DBset) CONNECT_ON_STARTUP=false, ~12 instances
Impact
Because the failure is fatal error rather than a panic, no recover() helps and the process exits (code 2). With CONNECT_ON_STARTUP=false every instance is left disconnected after the restart, and any instance without a stored JID needs a human to scan a QR again. A mutex around the map at whatsmeow.go:581 would stop a single instance pairing from taking down the whole server.
Happy to provide the full goroutine dump if useful.
- Ngôn ngữ chính
- Go
- Star
- 890
- Fork
- 473
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của evolution-foundation/evolution-go
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
evolution-foundation/evolution-go#193 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
evolution-foundation/evolution-go#104 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
evolution-foundation/evolution-go#101 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
evolution-foundation/evolution-go#97 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
evolution-foundation/evolution-go#204 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 5 ngày
Tất cả issue của evolution-foundation/evolution-go
Issue tương tự
-
area: global bug dx priority: low
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
grafana/mcp-grafana#1267 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
automation models
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
coverage-gap good-first-pattern help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
GoogleCloudPlatform/k8s-aibom#114 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
txn2/mcp-data-platform#1984 ·
Maintainer thường phản hồi trong vòng 1 ngày