Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Panic (nil pointer) in ReconnectClient kills the whole process - one instance takes down all others

Đang mở
#188 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 5 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
55/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
go
Lĩnh vực
backend

Hướng nghiên cứu

Đọc pkg/whatsmeow/service/whatsmeow.go, bắt đầu từ ReconnectClient quanh dòng 189 và handler Disconnected quanh các dòng 1985-1987. Tái hiện hoặc phân tích hai sự kiện đồng thời cho một instance, sau đó xác minh rằng các lần thử reconnect được tuần tự hóa hoặc loại bỏ trùng lặp một cách an toàn, và rằng một panic của reconnect không thể kết thúc process hoặc ảnh hưởng đến các instance khác.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

A race between two concurrent Disconnected events for the same instance
causes a nil pointer dereference in ReconnectClient. Because the panic happens
in an unrecovered goroutine, the entire Go process dies, disconnecting every
instance it hosts — not just the one that raced.

We run 7 instances in one process. Each panic takes all 7 offline.

Version
evoapicloud/evolution-go:0.7.2-beta
sha256:397d1da857616d073470c48aa02bdb1241e0e81c72cbec23a7f5af3411104b4d

Deployment: Docker Swarm, single replica, PostgreSQL 15 backend.

Stack trace
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x18 pc=0x112424e]

goroutine 471793 [running]:
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.whatsmeowService.ReconnectClient(...)
	/build/pkg/whatsmeow/service/whatsmeow.go:189 +0x24e
github.com/evolution-foundation/evolution-go/pkg/whatsmeow/service.(*MyClient).myEventHandler.func3({0xc00612f9b0, 0x24})
	/build/pkg/whatsmeow/service/whatsmeow.go:1987 +0xbb
created by ...myEventHandler in goroutine 471885
	/build/pkg/whatsmeow/service/whatsmeow.go:1985 +0x749f
The log right before the panic

Note the same instance firing the reconnect twice within the same second:

17:14:35 [INFO] [<id>] subscriptions [MESSAGE READ_RECEIPT] eventType Disconnected
17:14:35 [INFO] [<id>] Disconnected detected, restarting instance
17:14:35 [INFO] [<id>] Starting reconnection process - simulating restart
17:14:35 [INFO] [<id>] Disconnecting existing client
17:14:35 [ERROR] Error reading from websocket: failed to get reader: EOF
17:14:35 [INFO] [<id>] subscriptions [MESSAGE READ_RECEIPT] eventType Disconnected
17:14:35 [INFO] [<id>] Disconnected detected, restarting instance      <-- second one
17:14:35 [INFO] [<id>] Starting reconnection process - simulating restart
17:14:35 [INFO] [<id>] Disconnecting existing client
17:14:35 [INFO] [<id>] Cleaning up resources
panic: runtime error: invalid memory address or nil pointer dereference
Root cause (as we read it)

In myEventHandler, the *events.Disconnected case spawns a goroutine per event:

case *events.Disconnected:
    go func(instanceID string) {
        if err := mycli.service.ReconnectClient(instanceID); err != nil { ... }
    }(mycli.userID)

Two problems:

  1. No per-instance guard. Two Disconnected events for the same instance
    spawn two goroutines that both enter ReconnectClient. That function deletes
    from the shared maps (clientPointer, myClientPointer, killChannel)
    without a mutex, so one goroutine can remove what the other is about to
    dereference.

  2. No recover(). The goroutine is unprotected, so the panic is fatal to the
    process rather than to the reconnect attempt. Every other instance hosted in
    that process is disconnected as collateral.

Measured frequency

Single process, 5 connected instances, measured over 17 hours of uptime:

reconnect attempts 44 (≈1.76 per instance per hour)
crashes 2 in 48 hours
recovery time ~8 s (Swarm restart + all sessions reconnect)

Recovery is fast, but the blast radius is every instance in the process, and the
rate scales with the number of instances (more instances → more Disconnected
events → higher chance two collide).

Suggested fix

Two changes, independent of each other:

1. Serialize reconnects per instance (the actual fix)

// per-instance mutex, so a second Disconnected waits instead of racing
mu := w.reconnectMu(instanceID)
mu.Lock()
defer mu.Unlock()

Or drop the second event entirely if a reconnect for that instance is already in
flight — a sync.Map of in-flight instance IDs with LoadOrStore would do.

2. Recover in the goroutine (the safety net)

go func(instanceID string) {
    defer func() {
        if r := recover(); r != nil {
            logger.LogError("[%s] reconnect panicked, instance stays down: %v", instanceID, r)
        }
    }()
    ...
}(mycli.userID)

The second one alone would already change the outcome from "all instances die"
to "one instance fails to reconnect", which for a multi-tenant deployment is the
difference between an incident and a log line.

Happy to help

We can test a patched image against our workload (7 instances, ~6k messages/day)
and report back.


Confirmed in current source (2026-09-04)

We checked out tag 0.7.2 (commit 9337afc) and read
pkg/whatsmeow/service/whatsmeow.go. The handler is unchanged (line 1986):

// Trigger instance restart via websocket-capable service (non-blocking)
go func(instanceID string) {
    mycli.loggerWrapper.GetLogger(instanceID).LogInfo("[%s] Disconnected detected, restarting instance", instanceID)
    if err := mycli.service.ReconnectClient(instanceID); err != nil {
        mycli.loggerWrapper.GetLogger(instanceID).LogError("[%s] Failed to restart instance: %v", instanceID, err)
    }
}(mycli.userID)

In that whole 2,895-line file there is exactly one recover() (elsewhere,
unrelated) and one mutex (cachedWebVersionMu, unrelated). This goroutine has
neither.

We built and are running a patched image

To confirm the fix is viable we built 0.7.2 with only the recover() added, and
it is currently holding a live WhatsApp session in our staging process:

build           74 s, image builds clean from the public Dockerfile
gofmt           clean
startup         licensed, /instance/all returns 401 as expected
session         reconnected without QR ("Already logged in"), events subscribed

We are happy to open a PR with this change if you would like — it is 6 lines and
touches nothing else. The per-instance mutex (the actual race fix) would be a
separate, larger change and we would rather you decide the shape of that one.

Why this matters for multi-tenant deployments

The blast radius is every instance in the process, and the collision probability
grows with the number of instances hosted. For anyone running Evolution GO as a
shared backend, this effectively caps how many numbers can safely live in one
process.

Ngôn ngữ chính
Go
Star
890
Fork
473
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

  • Có Dockerfile hoặc tệp Docker Compose
  • Có mẫu pull request
  • Không có hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của evolution-foundation/evolution-go

Tất cả issue của evolution-foundation/evolution-go

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.