Deadlock: DynamicCodeGenerated allocates under _stubs_lock while the sampling signal handler spins in findRuntimeStub
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 52/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- cpp
- Lĩnh vực
- performance
Hướng nghiên cứu
Start with ddprof-lib/src/main/cpp/hotspot/jitCodeCache.cpp, focusing on DynamicCodeGenerated and findRuntimeStub; the issue describes the lock/allocation cycle and two possible fix directions. Reproduce or inspect the relevant locking and allocation paths, then add or run tests covering signal-handler lookup while the stub lock is held. Done means the deadlock is avoided without unbounded spinning; the issue names no specific test file.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
The native profiler (ddprof) deadlocks the JVM. A thread that holds the JitCodeCache stub lock calls malloc, while the profiler sampling signal interrupts another thread inside malloc, and that signal handler waits for the stub lock. All other sampled threads then spin in JitCodeCache::findRuntimeStub at 100% CPU, and the JVM never reaches a safepoint (jcmd cannot attach, the service stops answering).
We saw it twice in 4 hours on 2026-10-07, on 2 hosts, with the same cycle each time. We had seen the same symptom (N cores at 100%, JVM unresponsive, no log line) every few hours on this service before we captured stacks, on x86_64 and on arm64 (Graviton) hosts. Both captures below are from x86_64 hosts.
Versions
- dd-java-agent 1.65.1, which loads ddprof 1.49.0 (
libjavaProfiler) - Profiling on (
-Ddd.profiling.enabled=true), default profiler settings (CPU and wall-clock samplers) - Amazon Corretto 25.0.0+36 (JDK 25), x86_64, G1 GC
- Amazon Linux 2023, glibc malloc, kernel 6.1
- The application calls a native library (libvips, through glib) from many threads, so threads are often inside
malloc.
JitCodeCache::DynamicCodeGenerated and JitCodeCache::findRuntimeStub in ddprof-lib/src/main/cpp/hotspot/jitCodeCache.cpp are the same in v_1.49.0, v_1.51.0 and main (checked on 2026-10-07), so we think the newer versions are affected too.
The cycle
| Thread | Holds | Waits for |
|---|---|---|
JVM Service Thread, posting a deferred DynamicCodeGenerated event |
JitCodeCache::_stubs_lock (exclusive, taken in DynamicCodeGenerated) |
the glibc malloc lock (malloc from CodeCache::add) |
A worker thread interrupted by the profiler signal inside malloc (_int_malloc / malloc_consolidate) |
the glibc malloc lock | _stubs_lock: the signal handler calls findRuntimeStub → SpinLock::lockShared(), which spins with no bound |
| Every other sampled thread | — | _stubs_lock, spinning in lockShared() |
DynamicCodeGenerated calls _runtime_stubs.add() (which allocates) while it holds _stubs_lock, and findRuntimeStub takes the same lock from the signal handler with the unbounded lockShared(). Possible fixes: do not allocate under _stubs_lock (prepare the entry first), or use the bounded tryLockShared(max_spins) in findRuntimeStub when it runs in a signal handler and give up the frame.
Stacks, capture 1 (6 threads at 100% CPU, 5 of them in findRuntimeStub)
Collected with gdb -p <pid> -batch -ex "thread apply all bt". Addresses and paths removed.
Thread "Service Thread":
#0 __lll_lock_wait_private ()
#1 malloc ()
#2 CodeCache::add(void const*, int, char const*, bool) ()
#3 JitCodeCache::DynamicCodeGenerated(_jvmtiEnv*, char const*, void const*, int) ()
#4 JvmtiExport::post_dynamic_code_generated_internal(char const*, void const*, void const*) ()
#5 JvmtiDeferredEvent::post() ()
#6 ServiceThread::service_thread_entry(JavaThread*, JavaThread*) ()
#7 JavaThread::thread_main_inner() [clone .part.0] ()
#8 Thread::call_run() ()
#9 thread_native_entry(Thread*) ()
#10 start_thread ()
#11 clone3 ()
Thread "<HTTP worker thread>":
#0 JitCodeCache::findRuntimeStub(void const*) ()
#1 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, void const*, unsigned long, unsigned long, int, bool*) ()
#2 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, int, bool*) ()
#3 HotspotSupport::walkJavaStack(StackWalkRequest&) ()
#4 Profiler::recordSample(void*, unsigned long long, int, int, unsigned long long, Event*, unsigned long long*) ()
#5 CTimer::signalHandler(int, siginfo_t*, void*) ()
#6 <signal handler called> ()
#7 malloc_consolidate ()
#8 _int_malloc ()
#9 calloc ()
#10 g_malloc0 ()
#11 g_type_create_instance ()
#12 g_object_new_internal ()
#13 g_object_new_with_properties ()
...
Thread "<HTTP worker thread>":
#0 JitCodeCache::findRuntimeStub(void const*) ()
#1 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, void const*, unsigned long, unsigned long, int, bool*) ()
#2 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, int, bool*) ()
#3 HotspotSupport::walkJavaStack(StackWalkRequest&) ()
#4 Profiler::recordSample(void*, unsigned long long, int, int, unsigned long long, Event*, unsigned long long*) ()
#5 CTimer::signalHandler(int, siginfo_t*, void*) ()
#6 <signal handler called> ()
#7 hwy::x86::Cpuid (abcd=..., count=..., level=...)
#8 hwy::x86::FlagsFromCPUID ()
...
Stacks, capture 2 (8 threads at 100% CPU, all 8 in findRuntimeStub)
Thread "Service Thread":
#0 __lll_lock_wait_private ()
#1 malloc ()
#2 CodeCache::add(void const*, int, char const*, bool) ()
#3 JitCodeCache::DynamicCodeGenerated(_jvmtiEnv*, char const*, void const*, int) ()
#4 JvmtiExport::post_dynamic_code_generated_internal(char const*, void const*, void const*) ()
#5 JvmtiDeferredEvent::post() ()
#6 ServiceThread::service_thread_entry(JavaThread*, JavaThread*) ()
#7 JavaThread::thread_main_inner() [clone .part.0] ()
#8 Thread::call_run() ()
#9 thread_native_entry(Thread*) ()
#10 start_thread ()
#11 clone3 ()
Thread "<HTTP worker thread>":
#0 JitCodeCache::findRuntimeStub(void const*) ()
#1 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, void const*, unsigned long, unsigned long, int, bool*) ()
#2 HotspotSupport::walkVM(void*, _asgct_callframe*, int, StackWalkFeatures, EventType, int, bool*) ()
#3 HotspotSupport::walkJavaStack(StackWalkRequest&) ()
#4 Profiler::recordSample(void*, unsigned long long, int, int, unsigned long long, Event*, unsigned long long*) ()
#5 CTimer::signalHandler(int, siginfo_t*, void*) ()
#6 <signal handler called> ()
#7 _int_malloc ()
#8 calloc ()
#9 g_malloc0 ()
#10 g_closure_new_simple ()
#11 g_cclosure_new ()
#12 g_signal_connect_data ()
#13 ??? ()
...
top -H on both hosts showed the worker threads at 99.9% CPU each, while the VM Thread was sleeping (state S, under 4% CPU). jcmd <pid> Thread.print failed with AttachNotSupportedException ... doesn't respond within 10500ms.
We worked around it with DD_PROFILING_DDPROF_ENABLED=false (JFR-based profiling). We can test a fix build if useful.
- Ngôn ngữ chính
- C++
- Star
- 33
- Fork
- 14
- Merge trung bình
- 2 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 41
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của DataDog/java-profiler
-
needs-review upstream-tracking
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 15/100
DataDog/java-profiler#846 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
needs-review upstream-tracking
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
DataDog/java-profiler#829 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
needs-review upstream-tracking
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
DataDog/java-profiler#812 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
needs-review upstream-tracking
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
DataDog/java-profiler#785 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
needs-review upstream-tracking
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
DataDog/java-profiler#780 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của DataDog/java-profiler
Issue tương tự
-
[request] vsg/1.1.16Đang mởupstream update
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
conan-io/conan-center-index#31142 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
NVIDIA/DeepStream#78 ·