JVM SIGSEGV in vframe::java_sender via allocation-profiler JVMTI GetStackTrace on a virtual thread (JDK 25, Alpine/musl) — regression in 1.62.0 (1.60.1 OK)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- java
- Lĩnh vực
- observability-sre
Hướng nghiên cứu
Bắt đầu với hs_err_pid1.log đính kèm và truy vết đường đi của allocation profiler qua Profiler::recordJVMTISample, ObjectSampler::recordAllocation và JVMTI GetStackTrace trên các virtual thread. So sánh hành vi giữa các phiên bản tracer 1.60.1 và 1.62.0 trên JDK 25 với Alpine/musl; hoàn tất khi allocation profiling không còn tạo ra SIGSEGV đã được báo cáo.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Tracer Version(s)
1.62.0 (crashes). Reverting to 1.60.1 resolves it — so this is a regression introduced between 1.60.1 and 1.62.0.
Java Version(s)
OpenJDK Runtime Environment Temurin-25.0.1+8 (build 25.0.1+8-LTS), linux-amd64
JVM Vendor
Eclipse Adoptium / Temurin
OS / Environment
Alpine Linux v3.22 (musl libc), QEMU 4 cores / 11G, container -Xmx6g, G1 GC. Spring Boot 4.0.6 app, virtual threads in use. Continuous profiler enabled (native ddprof, allocation profiling on).
Bug Report
After upgrading the agent to 1.62.0, the JVM hard-crashes (SIGSEGV, is_crash) intermittently — ~10 min after startup under normal traffic. The crash is in the Datadog allocation profiler: on a sampled allocation it calls JVMTI GetStackTrace, and walking a virtual thread's frames segfaults inside vframe::java_sender().
This appears to be a sibling of #9830 but on a different stackwalk path: #9830 was the async/ASGCT path (vframeStreamForte::forte_next), mitigated by -Ddd.profiling.ddprof.cstack=vm (default since 1.55.0). This one is the JVMTI GetStackTrace path used by the allocation sampler on virtual threads, which cstack=vm does not appear to cover (we are on 1.62.0, well past that default, and still crashing).
si_code: 128 (SI_KERNEL), si_addr: 0x0; the faulting thread is a virtual-thread carrier (ForkJoinPool-1-worker), _thread_in_vm. The ArrayList.grow / QueryExecutorImpl.processResults Java frames are just the innocent allocation being sampled (no large allocation — verified, the underlying collection is small).
Workaround: downgrade to 1.60.1 (clean). We expect DD_PROFILING_ALLOCATION_ENABLED=false (or DD_PROFILING_DDPROF_ENABLED=false, given musl) would also avoid it.
Crash stack (top frames from hs_err_pid1.log):
# SIGSEGV (0xb) at pc=..., pid=1, tid=197
# Java VM: OpenJDK 64-Bit Server VM Temurin-25.0.1+8 (25.0.1+8-LTS, mixed mode, sharing, tiered, compressed oops, compressed class ptrs, g1 gc, linux-amd64)
# Problematic frame:
# V [libjvm.so+0x1190128] vframe::java_sender() const+0x38
Current thread: JavaThread "ForkJoinPool-1-worker-18" daemon [_thread_in_vm, id=197]
Native frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)
V [libjvm.so] vframe::java_sender() const+0x38
V [libjvm.so] JvmtiEnvBase::get_stack_trace(javaVFrame*, int, int, jvmtiFrameInfo*, int*)
V [libjvm.so] GetStackTraceClosure::do_vthread(Handle)
V [libjvm.so] JvmtiHandshake::execute(...)
V [libjvm.so] JvmtiEnv::GetStackTrace(...)
V [libjvm.so] jvmti_GetStackTrace
C [libjavaProfiler-dd-*.so] Profiler::recordJVMTISample(...)
C [libjavaProfiler-dd-*.so] ObjectSampler::recordAllocation(...)
C [libjavaProfiler-dd-*.so] ObjectSampler::SampledObjectAlloc(...)
V [libjvm.so] JvmtiExport::post_sampled_object_alloc(JavaThread*, oopDesc*)
V [libjvm.so] MemAllocator::allocate() / InstanceKlass::allocate_objArray(...)
V [libjvm.so] OptoRuntime::new_array_C(...)
Java frames: (J=compiled Java code, j=interpreted, Vv=VM code)
J java.util.ArrayList.grow() java.base@25.0.1
J org.postgresql.core.v3.QueryExecutorImpl.processResults(...)
J jdk.internal.vm.Continuation.run() java.base@25.0.1
J java.lang.VirtualThread.runContinuation() java.base@25.0.1
J java.util.concurrent.ForkJoinPool.runWorker(...) java.base@25.0.1
j java.util.concurrent.ForkJoinWorkerThread.run() java.base@25.0.1
siginfo: si_signo: 11 (SIGSEGV), si_code: 128 (SI_KERNEL), si_addr: 0x0
I can attach the full hs_err_pid1.log (and the .jfr) if useful.
- Ngôn ngữ chính
- Java
- Star
- 737
- Fork
- 361
- Merge trung bình
- 3 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 173
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của DataDog/dd-trace-java
-
type: feature request
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 70/100
DataDog/dd-trace-java#10245 · 1 bình luận ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 62/100
DataDog/dd-trace-java#12608 ·
-
type: bug report
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
DataDog/dd-trace-java#12597 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
DataDog/dd-trace-java#12540 · 3 bình luận · 1 người được giao ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 25/100
DataDog/dd-trace-java#12480 ·
Tất cả issue của DataDog/dd-trace-java
Issue tương tự
-
certification
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's Đang mởbug ecr
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Needs: Triage Type: Feature request
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
AntennaPod/AntennaPod#8794 ·
-
agentic-workflows
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
github/copilot-sdk#2760 ·