Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

🐛[BUG]: cdf in metrics.general.histogram drops earlier inputs from the last cumulative bin

Aperta
#2,005 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
35/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python, pytorch

Direzione di ricerca

Inizia da physicsnemo/metrics/general/histogram.py, concentrandoti su _high_memory_bin_reduction_cdf, _low_memory_bin_reduction_cdf, _count_bins e Histogram.update. Leggi test/metrics/test_metrics_general.py ed esegui i test dell’istogramma; aggiungi una copertura di regressione separata per gli input cdf cumulativi e per il conteggio dei bin dopo update. Il lavoro è completato quando cdf(x, y) corrisponde a cdf(torch.cat((x, y))) e number_of_bins è uguale a bin_edges.shape[0] - 1.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Version

main at ff5d19d08123de47ca446caed1d70a225d540184

On which installation method(s) does this occur?

Source

Describe the issue

_high_memory_bin_reduction_cdf in physicsnemo/metrics/general/histogram.py writes the size of the current input into the last cumulative bin instead of adding it. Its low memory twin _low_memory_bin_reduction_cdf adds it. _count_bins always tries the high memory routine first and only falls back on a RuntimeError, so whenever counts are accumulated across calls the last bin ends up holding only the last input.

The public cdf function hits this as soon as it gets more than one input. _compute_counts_cdf loops over the inputs and feeds the running counts back into _count_bins, so after cdf(x, y, bins=10) the last bin holds len(y) while every other bin holds counts over x and y together. The normalization counts / counts[-1] then divides by the wrong total and the lower bins come out above one. With 10 samples in x and 5 in y the largest value is 2.8 instead of 1.0.

A second bookkeeping problem sits in Histogram.update. It stores self.bin_edges.shape[0] as number_of_bins, which is the number of edges. __init__ uses bins.shape[0] - 1. After an update that extends the bin range, number_of_bins is one larger than the number of rows in counts, and a later __call__ builds the histogram with one bin more than before.

The existing test_histogram in test/metrics/test_metrics_general.py compares the low and high memory routines, but it passes the same counts tensor object to both. Each routine mutates and returns that tensor, so the comparison is between one tensor and itself and cannot catch the first problem.

I expected cdf(x, y) to equal cdf(torch.cat((x, y))) and Histogram.number_of_bins to stay equal to bin_edges.shape[0] - 1 after update.

Minimum reproducible example
import torch
import physicsnemo.metrics.general.histogram as hist

torch.manual_seed(0)
x = torch.randn(10, 3, 4)
y = torch.randn(5, 3, 4)
bin_edges, cdf = hist.cdf(x, y, bins=10)
print(cdf.max().item())
# 2.799999952316284 on main, has to be 1.0

H = hist.Histogram((1, 3, 4), bins=10)
H(x)
bin_edges, counts = H.update(x + 10.0)
print(H.number_of_bins, bin_edges.shape[0] - 1, counts.shape[0])
# 92 91 91 on main, all three have to agree
Relevant log output
CDF_MAX 2.799999952316284
NBINS 92 91 91
Environment details
Bare-metal, CPU only, Python 3.12, torch 2.14.0+cpu, physicsnemo installed from source with pip install -e .

I have a two line fix ready with CPU regression tests for both problems, and I will open a PR that references this issue.

Lingua principale
Python
Stelle
3.3k
Fork
787
Merge medio
3g 4h
PR unite (30g)
28

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/physicsnemo

Tutte le issue di NVIDIA/physicsnemo

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.