A heartbeat that blocks forever is never detected, unlike one that fails
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 42/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- postgresql, ruby
- Lĩnh vực
- backend, databases, distributed-systems
Hướng nghiên cứu
Start by reading SolidQueue::Process#heartbeat and launch_heartbeat, then inspect Supervisor::Maintenance#launch_maintenance_task and supervise. Compare how blocked heartbeat and maintenance tasks behave with the current termination checks. Done requires an agreed approach that detects or prevents indefinitely blocked work, including the supervisor behavior or documented database socket configuration described in the issue.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
(Written by Claude on @julik's behalf)
Follow-up to #763, which I filed and which was closed as completed by #778. #778 fixes the case I could reproduce at the time; this is the neighbouring case, which I have now hit in production and which I believe is still open.
The distinction
#778 stops a process whose heartbeats keep failing:
def heartbeat
process&.heartbeat
rescue ActiveRecord::RecordNotFound
stop_to_be_replaced
rescue => error
stop_to_be_replaced if presumed_dead?
raise error
end
Both new paths hang off a raise. But a heartbeat can also never return at all, and then neither branch is reachable.
SolidQueue::Process#heartbeat does a real round-trip:
restore_attributes
with_lock { touch(:last_heartbeat_at) }
If the socket under that connection is half-open - the server process is alive enough for its kernel to keep the connection, but nothing is being served - the thread blocks in the driver indefinitely. No exception is raised, so presumed_dead? is never evaluated.
Why the timer does not save you
launch_heartbeat runs the heartbeat inside a Concurrent::TimerTask. In concurrent-ruby (checked against 1.3.7), execute_task is:
_success, value, reason = @task.execute(self)
if completion.try?
self.value = value
schedule_next_task(calculate_next_interval(start_time))
...
observers.notify_observers { [time, self.value, reason] }
The reschedule and the observer notification both happen after the task returns. A run that never returns is never rescheduled and never reported - the heartbeat thread goes quiet permanently, having neither succeeded nor failed.
There is also no configuration escape: TimerTask#timeout_interval is now a no-op that warns "TimerTask timeouts are now ignored as these were not able to be implemented correctly".
The supervisor has the same property
This is the part I think is most worth looking at. Supervisor::Maintenance#launch_maintenance_task is also a Concurrent::TimerTask, running prune_dead_processes. So when a supervisor's maintenance run blocks on the same unresponsive database, pruning stops too - the mechanism meant to notice dead processes is built from the same material that just died.
Meanwhile supervise only calls check_and_replace_terminated_processes, so a forked child that blocks without exiting is never replaced. Between the three, a process can be alive, registered, holding claimed executions, and doing nothing whatsoever, with nothing in the system positioned to notice.
What it looked like in production
A PgBouncer instance in front of our Postgres froze without dying - process alive, event loop wedged, kernel holding every TCP connection half-open. Our scheduler's heartbeat thread blocked mid-with_lock and stayed there. The recurring scheduler runs in exactly one process, so all 33 recurring tasks stopped for 101 minutes and only came back on a manual redeploy. The supervisor never replaced the child; other supervisors pruned the process row but pruning sends no signal, so it kept claiming jobs and those claims later failed as ProcessMissingError.
Left alone, the sockets would have cleared on Linux's default 7200s keepalive - about 2h11m.
What actually fixed it for us, and why I am not sure it is your problem
Setting client-side socket bounds in database.yml:
keepalives: 1
keepalives_idle: 30
keepalives_interval: 10
keepalives_count: 3
tcp_user_timeout: 60000
connect_timeout: 5
tcp_user_timeout is the load-bearing one - TCP keepalive only probes a socket with no unacknowledged data in flight, which is the wrong half of the problem. With it, the block becomes a raise after ~60s, and #778's logic then engages exactly as designed.
So one defensible answer is "this is a libpq configuration problem, document it and close". I would understand that. What makes me file anyway:
- The gap is invisible. An app with default
database.ymlhas no socket bound, and nothing in SolidQueue's own supervision will notice, so the failure mode is a silent stop rather than a loud one. - It is not Postgres-specific. Any adapter whose socket can go quiet has the same shape.
- The supervisor's maintenance task sharing the defect seems worth fixing regardless of what the app does with its sockets.
Possible directions
Not a proposal, just what seems available:
- Track the last time a heartbeat completed in memory, separately from
last_heartbeat_at, and have the run loop stop the process when that goes older thanprocess_alive_threshold. This is the only option that does not depend on the blocked thread itself. - Have supervisors act on prunable children rather than only on exited ones - prune currently deletes the row and sends nothing.
- At minimum, README guidance on socket-level timeouts for the database connection, since without one nothing else here can engage.
Happy to have a go at (1) if you think it is the right shape.
Refs
- #763 - Scheduler becomes an un-replaced zombie after host sleep/resume (mine, closed by #778)
- #778 - Stop processes whose heartbeats keep failing past the alive threshold (merged 2026-08-22, after v1.7.0 was cut, unreleased at time of writing)
- #751 - Supervisor remains alive with missing worker process; due scheduled jobs stop being dispatched
- #760 - Replace forked processes that fail to boot (covers stalling during startup, not after)
- #781 / #791 - Replace terminated forks/threads even if releasing claimed jobs fails
- #788 - Retry finalizing claimed executions on transient errors (open)
- #716 - the earlier nil-process heartbeat fix
Versions: solid_queue 1.6.0, concurrent-ruby 1.3.7, Ruby 3.4.2, PostgreSQL via PgBouncer in transaction mode.
- Ngôn ngữ chính
- Ruby
- Star
- 2.5k
- Fork
- 252
- Merge trung bình
- 8 giờ 31 phút
- Pull request đã merge (30 ngày)
- 5
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Có Dockerfile hoặc tệp Docker Compose
- Không có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của rails/solid_queue
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
rails/solid_queue#806 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
rails/solid_queue#805 ·
-
limits_concurrency on_conflict: :discard looking only into running jobs and not blocked jobsĐang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 52/100
rails/solid_queue#804 · 1 bình luận ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 35/100
rails/solid_queue#802 ·
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
rails/solid_queue#797 ·
Tất cả issue của rails/solid_queue
Issue tương tự
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
jetrockets/jet_ui#47 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
FreeCAD/homebrew-freecad#870 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Messages with a markdown image that has a relative or malformed URL throw TypeError: Invalid URLĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày