Queueing-time profiler aborts the whole instrumentation install under a JDK 24+ AOT cache (zero spans); disabling that one feature is enough
メンテナーはふだん 2 日以内に返信
@mcculls がすでに取り組んでいます。
2026年9月21日 から。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
- 技術スタック
- java
調査の方向性
Start with UnwrappingVisitor and the retransformation path involving RedefinitionStrategy and BatchAllocator. Reproduce with a JDK 25 AOT cache and the queueing-time profiler enabled, then verify that a failed optional transformation no longer aborts installation, spans remain available, and the degradation is logged at info level.
索引モデルが issue の本文から書いたものです。
説明
Tracer Version(s)
1.65.1, 1.66.0
Java Version(s)
25
JVM Vendor
Eclipse Adoptium / Temurin
Bug Report
TL;DR: with a JDK 25 AOT cache, the queueing-time profiler's failed retransformation batch silently aborts the entire instrumentation install (zero spans). Disabling that single feature — -Ddd.profiling.queueing.time.enabled=false — restores all spans and keeps every other profiling capability. Fully disabling the profiler is not necessary.
Behavior
With -XX:AOTCache enabled (jib containerizingMode=packaged, Spring Boot, JDK 25):
- The tracer starts normally, prints its configuration, reports
agent_error: false. - Not a single span is produced. No warnings at default log level.
- Only
-Ddd.trace.debug=truereveals:Exception while retransforming 574 classes.
Identical behaviour on 1.65.1 and 1.66.0, so not a recent regression.
Root cause
UnwrappingVisitor (queueing-time profiling) adds the TaskWrapper interface to the classes it instruments. Adding an interface is a structural change, which the JVM rejects on retransformation. Normally this works because the agent installs before those classes load — but a JDK 24+ AOT cache (JEP 483) materializes them before premain.
Since no BatchAllocator is configured, that is the single batch holding every class, and RedefinitionStrategy aborts the whole install. One unusable profiler feature takes all instrumentation down with it.
Why this is filed separately from #10479
#10479 covers partial span loss via the context-store weak-map fallback (addressed by #12105). This is a different mechanism with a total, silent loss — and a much narrower fix surface.
Expected Behavior
A retransformation failure of one optional profiling feature should not abort the entire instrumentation install, and the degradation should be visible at default log level.
Workaround
-Ddd.profiling.queueing.time.enabled=false — all spans return, all other profiling features (CPU, allocation, lock, IO, …) keep working.
The failing setup is exactly the one from the docs, with profiling enabled. Training with the explicit -javaagent:dd-java-agent.jar=aot_training argument fails the same way.
-XX:-AOTClassLinking also avoids it, but gives up the startup gain entirely: on our Spring service startup was 20.5s without a cache, 12.8s with it, and 21.0s with -XX:-AOTClassLinking.
Fix
#12506 retries the failed batch without the interface change. Queueing-time profiling itself keeps working — TaskWrapper is only read by QueueTimeEvent.setTask() and is guarded by instanceof, so the events, durations, scheduler, queue type and span ids are all still recorded; only the task field loses its type resolution (wrapper class instead of the real one). JFR table in #12506. The degradation is logged once at info. With that fix the workaround flag is no longer needed.
Reproduction Code
- Spring Boot app, jib
containerizingMode=packaged, JDK 25 (Temurin). - Training run:
-XX:AOTCacheOutput=app.aotwith the agent attached, then run with-XX:AOTCache=app.aot -javaagent:dd-java-agent.jar. - Send requests: zero spans. Add
-Ddd.profiling.queueing.time.enabled=false: all spans appear (245 spans in our 20-request measurement, identical to no-AOT baseline — table in #12506).
- 主要言語
- Java
- スター
- 737
- フォーク
- 362
- 平均マージ
- 3日 18時間
- マージ済み PR(30日)
- 180
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
DataDog/dd-trace-java のほかの issue
-
Update DEFAULT_HTTP_CLIENT_ERROR_STATUSES to include 5xx status codes再び着手できるかも このイシューのプルリクエストはマージされずにクローズされました。 オープンtype: feature request
難易度 1/5 1〜3時間 初心者へのやさしさ 70/100
DataDog/dd-trace-java#10245 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
gRPC server instrumentations (grpc-1.5, armeria-grpc) don't make extracted W3C baggage current in the handler対応中かも @mcculls が 7 日前に担当しました。 オープンcomp: context propagation inst: grpc type: feature request
難易度 3/5 1〜2日 初心者へのやさしさ 78/100
DataDog/dd-trace-java#12654 · 担当者 1 名 ·
メンテナーはふだん 2 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 62/100
DataDog/dd-trace-java#12608 ·
メンテナーはふだん 2 日以内に返信
-
type: bug report
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
DataDog/dd-trace-java#12597 ·
メンテナーはふだん 2 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 25/100
DataDog/dd-trace-java#12480 ·
メンテナーはふだん 2 日以内に返信
DataDog/dd-trace-java の issue をすべて見る
似ている issue
-
[destination-snowflake] Custom domains rejected unlike source connections対応中かも @kuza55 が今日担当しました。 オープンautoteam community connectors/destination/snowflake team/use
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
area-dashboard
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
component/operate kind/feature-request
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
Forge coverage prompts carry text the agent cannot act on対応中かも @graalvmbot が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
oracle/graalvm-reachability-metadata#10572 ·
メンテナーはふだん 1 日以内に返信
-
[CI] Core CI doesn't run for changes to amoro-format-lance (and amoro-web)対応中かも @MarkAlex1234 が今日担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信