Classic-only histogram: consider synchronized block instead of multi-LongAdder for observe() hot path
@zeitlinger is already working on this.
Since Feb 25, 2026.
Assessment
This issue has not been assessed yet.
Description
Context
While benchmarking the Prometheus shim PoC (bridging Prometheus client API to the OTel SDK), I found that classic-only histograms are 30% faster through the OTel SDK than through native Prometheus.
Benchmark numbers (JMH, single thread)
| Path | observe() latency |
|---|---|
| Native Prometheus (classic-only) | 10.5 ns |
| OTel SDK (explicit bucket histogram) | 7.3 ns |
Root cause
Native Prometheus doObserve() uses 3 separate CAS-based atomics per call:
classicBuckets[i].add(1)—LongAddersum.add(value)—DoubleAddercount.increment()—LongAdder
Plus a buffer.append() CAS attempt and volatile reads for reset/scale-down state.
The OTel SDK uses a single synchronized block with plain +=/++ arithmetic:
synchronized (lock) {
this.sum += value;
this.count++;
this.counts[bucketIndex]++;
// min/max tracking
}
In uncontended (single-thread) benchmarks, HotSpot elides the uncontended lock and optimizes the plain arithmetic freely, beating the multi-CAS approach.
Suggestion
For classic-only histograms (where nativeInitialSchema == CLASSIC_HISTOGRAM), consider an alternative doObserve() implementation that uses a synchronized block with plain fields instead of multiple LongAdder/DoubleAdder instances. The buffer mechanism (needed for native histogram scale-down) could also be bypassed in classic-only mode.
This wouldn't affect native or hybrid histograms, which still need the current design.
Multi-threaded consideration
The LongAdder approach was chosen for multi-threaded scalability (striped cells reduce contention). A synchronized block would serialize threads. However:
- Most real-world
observe()calls happen on different label-value combinations (different data points), so contention on a single data point is rare - Even under contention, the critical section is very short (~5 ns of arithmetic), so lock hold time is minimal
- A benchmark with 4 threads would clarify the actual tradeoff
Not a high priority — 10.5 ns is already excellent. But worth considering if classic histogram performance matters.
- Dominant language
- Java
- Stars
- 2.3k
- Forks
- 833
- Avg merge
- 2d 16h
- Merged PRs (30d)
- 86
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from prometheus/client_java
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
prometheus/client_java#2416 · 1 comment ·
-
Switch Micrometer compatibility workflow to upstream once typed-descriptor path becomes default Open
Difficulty 1/5 Under an hour Newbie friendliness 86/100
prometheus/client_java#2182 ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
prometheus/client_java#2306 · 9 comments · 4 reactions ·
-
Difficulty 5/5 Over a week Newbie friendliness 32/100
prometheus/client_java#2084 · 3 comments ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
prometheus/client_java#2075 · 10 comments · 1 reaction ·
All issues in prometheus/client_java
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
infinispan/infinispan#18150 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
untriaged
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
opensearch-project/k-NN#3597 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100