Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

test_relpos_attention_local: local (rel_pos_local_attn) attention diverges from NeMo

Aperta
#44 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
52/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
cpp, python

Direzione di ricerca

Inizia con forward_local e l'eseguibile test_relpos_attention_local, quindi confronta la sua gestione del layout pos_emb scaricato e della finestra di attenzione con scripts/gen_nemo_baseline.py. Riproduci la divergenza CPU f32 a W=64 e W=32; il lavoro è completato quando l'output C++ risponde a W e corrisponde alla baseline NeMo, mentre i test chunked e memory continuano a passare.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

While building a full NeMo-baseline set to run the model-dependent test suite, I
hit a divergence in the local (Longformer) attention path that looks separate
from #39 (the streaming O(N²) fix) — filing it on its own.

Symptom

test_relpos_attention_local fails on the 110m anchor, on CPU (PARAKEET_DEVICE=cpu,
f32 GGUF), so it isn't iGPU fp16 tolerance:

[relpos_attention_local] n=47616 max|d|=3.349e+02 mean|d|=9.779e+00 (worst@47338 got=0.44526 ref=335.37750) -> FAIL

The divergence is broad (mean |d| ≈ 10, not a single element) and the worst point
is the last time frame (worst index 47338 = frame 92 of T=93, d_model=512).

It's not the --att-context-size (W) chosen for the baseline

I regenerated PARAKEET_TEST_BASELINE_LOCAL at two windows and re-ran:

W result
64 worst@47338 got=0.44526 ref=335.37750
32 worst@47338 got=0.44526 ref=356.31686

The C++ output (got) is identical across W while NeMo's ref changes — i.e.
forward_local does not respond to the window the baseline encodes. (W=128 is
correctly rejected by the test since W ≥ T.)

test_relpos_attention_local_chunked and test_relpos_attention_local_memory
pass (they use an internal brute-force reference), so the gap is specific to
the non-chunked forward_local vs the NeMo rel_pos_local_attn baseline.

Reproduce

# baseline (NeMo): local attention with a finite window over speech.wav
python scripts/gen_nemo_baseline.py \
  --model nvidia/parakeet-tdt_ctc-110m \
  --audio tests/fixtures/speech.wav \
  --att-context-size 64 --output /tmp/baseline_local.gguf

# convert the 110m anchor to f32 gguf -> PARAKEET_TEST_GGUF
PARAKEET_DEVICE=cpu \
PARAKEET_TEST_GGUF=/tmp/pk110m-f32.gguf \
PARAKEET_TEST_BASELINE_LOCAL=/tmp/baseline_local.gguf \
  ./build/tests/test_relpos_attention_local

Question

Is this a known limitation, a layout/convention mismatch between the dumped
pos_emb ([2W+1, d_model]) and what forward_local expects, or a real bug in
the non-chunked local path? Happy to dig into forward_local if it's worth a fix.

Lingua principale
C++
Stelle
786
Fork
93
Merge medio
9g 19h
PR unite (30g)
4

Preparare l'ambiente

Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/parakeet.cpp

Tutte le issue di mudler/parakeet.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.