Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

0.7.2: fatal error: concurrent map writes in whatsmeowService.StartClient (kills the whole process)

Đang mở
#203 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 5 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
62/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
go
Lĩnh vực
authentication, backend

Hướng nghiên cứu

Start in pkg/whatsmeow/service/whatsmeow.go at the map deletion on line 581, then trace the recursive StartClient retry at line 629 and StartInstance at line 2407. Reproduce concurrent /instance/connect calls while pairing by QR, with handleQRCodes active. Done means the race no longer produces a fatal concurrent map write or process-wide exit, and QR pairing can complete.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Running evoapicloud/evolution-go:0.7.2 with ~12 instances, the process dies repeatedly with an unrecoverable Go runtime error. Not a recoverable panic — it takes down every instance in the deployment. We are at 168 restarts in 54 days.

Stack trace

fatal error: concurrent map writes

goroutine 18577 [running]:
internal/runtime/maps.fatal({0x195b5e4?, 0x0?})
	/usr/local/go/src/runtime/panic.go:1046 +0x18
internal/runtime/maps.(*Map).Delete(0xc000ffe9c0, 0x171eb60, 0xc003c3b228)
	/usr/local/go/src/internal/runtime/maps/map.go:652 +0x4c
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
	/build/pkg/whatsmeow/service/whatsmeow.go:581 +0x28bc
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
	/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
	/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartClient(...)
	/build/pkg/whatsmeow/service/whatsmeow.go:629 +0x3185
created by github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.StartInstance in goroutine 18558
	/build/pkg/whatsmeow/service/whatsmeow.go:2407 +0xa2d

The crashing write is a map delete at whatsmeow.go:581, reached from StartClient. Note StartClient appears recursively via whatsmeow.go:629 — it re-enters itself on retry, so a single instance can have several StartClient frames in flight while other instances are starting from StartInstance in their own goroutines. The map mutated at :581 looks like shared state across instances with no mutex.

When it happens

Every crash we have captured has a live handleQRCodes goroutine in the same dump:

goroutine 72775 [sleep]:
time.Sleep(0xdf8475800)
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.(*MyClient).handleQRCodes.func1()
	/build/pkg/whatsmeow/service/whatsmeow.go:785 +0x85
created by github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.(*MyClient).handleQRCodes in goroutine 72773
	/build/pkg/whatsmeow/service/whatsmeow.go:704 +0x86

So the trigger for us is: pair an instance by QR while another instance is (re)connecting. The pairing never reaches the auth store — after the restart the instance has an empty JID and needs a new QR. From the outside it looks like "we scanned the QR and it dropped by itself".

Reproduction

  1. Have several instances registered, some offline.
  2. Call POST /instance/connect on more than one of them at roughly the same time (an external reconnect loop is enough).
  3. Pair one instance by QR while that is happening.

It is a race, so it does not reproduce every time — for us it lands about 3 times a day.

Environment

  • evoapicloud/evolution-go:0.7.2 (sha256:6fa601464bd76d3d19ece4c48d8bb5373d025d95d45b939ae4bf0c77b03f5aaa)
  • whatsmeow v0.0.0-20260630180629-b572e5bcb92b
  • Auth store in Postgres (POSTGRES_AUTH_DB set)
  • CONNECT_ON_STARTUP=false, ~12 instances

Impact

Because the failure is fatal error rather than a panic, no recover() helps and the process exits (code 2). With CONNECT_ON_STARTUP=false every instance is left disconnected after the restart, and any instance without a stored JID needs a human to scan a QR again. A mutex around the map at whatsmeow.go:581 would stop a single instance pairing from taking down the whole server.

Happy to provide the full goroutine dump if useful.

Ngôn ngữ chính
Go
Star
890
Fork
473
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

  • Có Dockerfile hoặc tệp Docker Compose
  • Có mẫu pull request
  • Không có hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của evolution-foundation/evolution-go

Tất cả issue của evolution-foundation/evolution-go

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.