Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

I2_S produces incorrect output on aarch64

Offen
#598 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
25/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Veraltet
Tech-Stack
cpp

Rechercherichtung

Beginne im angehefteten 3rdparty/llama.cpp-Submodul und konzentriere dich auf quants.c, ggml-cpu-i2s.c, ggml_gemm_i2_i8_s und ggml_vec_dot_i2_i8_s. Reproduziere das Problem mit dem aufgeführten llama-cli-Befehl auf aarch64 und vergleiche anschließend Perplexity sowie die Ausgabe des GEMV/GEMM-Kernels mit x86. Als erledigt gilt die Aufgabe, wenn Ausgabe und Perplexity architekturübergreifend übereinstimmen.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

I2_S inference is broken on aarch64. The model loads and runs, but output is nonsense — it does not crash, so it looks like a bad model rather than a broken kernel.

On an Orange Pi 5 Plus (RK3588, Debian 12, GCC 12.2) with microsoft/BitNet-b1.58-2B-4T's official ggml-model-i2_s.gguf:

$ llama-cli -m ggml-model-i2_s.gguf -p "The capital of France is" -n 12 --temp 0
> The capital of France is  ????????????????
[ Prompt: 0.7 t/s | Generation: 0.7 t/s ]

The same file on x86-64 gives The capital of France is Paris. at ~40 t/s. Model md5 verified identical on both machines.

Three separate bugs, all in 3rdparty/llama.cpp code paths that x86 never compiles:

  1. QK_I2_S is 128 under AVX2 but 64 under __ARM_NEON, in both quants.c and ggml-cpu-i2s.c. It is the on-disk block size, so it must match the file format on every architecture.
  2. The scalar vec_dot fallback decodes the block-interleaved weight layout sequentially.
  3. ggml_gemm_i2_i8_s's ACT_PARALLEL branch inverts ggml_vec_dot_i2_i8_s's nrc semantics, corrupting prefill.

There is also no NEON path for I2_S at all — aarch64 unpacks one 2-bit weight at a time.

Fixes in https://github.com/isHuangXin/llama.cpp/pull/2, against the pinned 3rdparty/llama.cpp submodule. After them, perplexity on aarch64 matches x86 to 0.161% (74.0952 vs 73.9758, same model and corpus), and kernel output is bit-identical for both GEMV and GEMM.

This may be the same root cause as #55.

Vorherrschende Sprache
C++
Sterne
40.3k
Forks
3.7k
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Entwicklungsumgebung

Dieses Projekt bietet weder Dev-Container noch Dockerfile noch Beitragsleitfaden – die Einrichtung liegt bei Ihnen. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus microsoft/BitNet

Alle Issues in microsoft/BitNet

Ähnliche Issues

Weitere Issues zu C++

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.