Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

Latent int16 overflow in ggml_vec_dot_i2_i8_s_1x1: 128 maddubs results per lane against a 32767 ceiling

Offen
#629 5 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
56/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
c, cpp
Bereich
backend, performance

Rechercherichtung

Beginne in ggml/src/ggml-cpu/quants.c bei ggml_vec_dot_i2_i8_s_1x1 und verfolge die int16-Akkumulation durch die 32-Iterations-Schleife und den abschließenden Fold. Reproduziere das Problem mit den BitNet-b1.58-2B-4T-Gewichten und +127 int8-Aktivierungen und vergleiche anschließend die Fallback-Ausgabe mit exakter int64-Arithmetik. Erledigt ist dies, wenn der Fallback nicht mehr überläuft und weiterhin korrekte Ergebnisse für gewöhnliche Aktivierungen liefert.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Summary

ggml_vec_dot_i2_i8_s_1x1 accumulates 128 vpmaddubsw results into one
int16 lane before folding. The worst case is 128 * 508 = 65024 against a
32767 ceiling. It is correct in practice only because signed activations
cancel, and it is latent on a default build because the sgemm path takes
the work instead.

I am filing it as hardening, not as a wrong-output bug, and the distinction is
deliberate — see Severity below.

The arithmetic

ggml/src/ggml-cpu/quants.c:

const int group32_num = nb / 32;                     // :1396
...
__m256i accu32 = _mm256_setzero_si256();             // :1413  (int16 lanes)
for (int j = 0; j < 32; j++) {                       // :1414
    ...
    accu32 = _mm256_add_epi16(accu32, _mm256_add_epi16(xq8_0, xq8_1));   // :1439
    accu32 = _mm256_add_epi16(accu32, _mm256_add_epi16(xq8_2, xq8_3));   // :1440
}
accu = _mm256_add_epi32(_mm256_madd_epi16(accu32, one16), accu);         // :1445

32 iterations × 4 planes = 128 vpmaddubsw results per int16 lane. Each is at
most 2 * 2 * 127 = 508, so the lane can reach 65024, roughly twice the
ceiling. The name accu32 suggests 32-bit; the type is __m256i used with
_mm256_add_epi16.

What actually happens

With all activations the same sign the sum does not cancel and the lane wraps.
Measured on the real weights of BitNet-b1.58-2B-4T — 30 tensors at
K = 6912, driven against int8 activations pinned to +127 — 11,998 of
12,000 rows
disagree with exact int64 arithmetic, and every error is a
multiple of 65536, the signature of a wrapped int16.

With ordinary int8 activations, 0 of 12,000 rows wrap. That is why it has
never been seen.

Severity: low, and I want to be accurate about why

  • On a build where llamafile_sgemm handles these shapes this function is not
    called at all. The kernel's own counters on my build read
    calls=0 sgemm=52068.
  • It becomes reachable when that path is disabled, which is a supported switch
    rather than an exotic one.
  • Even then it needs same-signed activations across a whole group of 32 blocks.

So: a latent overflow in a fallback, not a correctness bug users are hitting.
I would not have filed it on its own; I am filing it because it is the same
missing bound-derivation as the fold issue I opened alongside this one (#628), in the
same file family, and a fix for one is naturally a fix for both.

Not a duplicate of the ARM reports

#411 and the audit in #525 report an int16 overflow, but it is the NEON
vmlal_s8 accumulation in src/ggml-bitnet-mad.cpp, which does produce visible
garbage. This is a different function, a different instruction set and a
different failure mode — this one is invisible precisely because it cancels.

Suggested fix

Fold every 8 groups rather than every 32, or accumulate in int32 directly.
Upstream ggml's equivalent kernels derive the bound explicitly and stay inside
it — TQ2_0 in arch/x86/quants.c carries // 16-bit sums, because 256*127 still fits — which is the pattern worth copying here.

Reproducible from https://github.com/purpleskulll/bitnet-t5b.

Vorherrschende Sprache
C++
Sterne
40.3k
Forks
3.7k
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Entwicklungsumgebung

Dieses Projekt bietet weder Dev-Container noch Dockerfile noch Beitragsleitfaden – die Einrichtung liegt bei Ihnen. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus microsoft/BitNet

Alle Issues in microsoft/BitNet

Ähnliche Issues

Weitere Issues zu C++

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.