[Bug]: I2_S GEMM fast path produces garbage for multi-token prompts on AVX-only CPUs
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 68/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- cmake, cpp
调研方向
从 ggml-cpu.c 中 1492 行附近的 I2_S GEMM fast path 开始,检查 ggml_gemm_i2_i8_s 在 nr e 1 时的列 stride 处理。使用文档中说明的 CMake flags 在禁用 AVX2 的情况下进行构建,然后使用较短和较长的 prompt 重现问题,并将 GEMM path 与 1217 行附近正常工作的 GEMV path 进行比较。当多 token 的 I2_S prompt 在 AVX-only CPU 上能够生成连贯输出,且无需禁用 fast path 时,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Description
On CPUs without AVX2 (e.g., Intel Xeon E5-2690 v2, Ivy Bridge, AVX-only), the I2_S GEMM fast path in ggml-cpu.c:1492 produces corrupt output when processing prompts with more than ~3 tokens. Single-token generation (GEMV path) works correctly.
Symptoms
- 1-3 token prompts → coherent output
- 5+ token prompts →
??????or garbled output - The corruption affects the KV cache: even after the prompt is processed, subsequent generation tokens are garbled
Root Cause
The GEMM fast path (ggml_gemm_i2_i8_s) is called for multi-token prompt evaluation (when src1 has multiple columns). The scalar fallback implementation has a bug in how it indexes the I2_S weight matrix and/or activation matrix for column strides > 1.
Workaround
Disabling the GEMM fast path forces I2_S through the dequantize-then-float-matmul path:
// ggml-cpu.c:1492 — change from:
if (src0->type == GGML_TYPE_I2_S && ggml_n_dims(src0) == 2) {
// to:
if (false && src0->type == GGML_TYPE_I2_S && ggml_n_dims(src0) == 2) {
This produces correct results but is slower (~0.6 tok/s prompt eval vs ~26 tok/s for F16 on the same hardware).
Key Distinction from #547
This is distinct from #547/PR #580 which covers the empty-body scalar fallback for ggml_vec_dot_i2_i8_s_* kernels. After applying those fixes (or equivalent scalar implementations), the GEMM path still produces wrong results for multi-token prompts.
The GEMV path (line ~1217) works correctly for single-token generation. The issue is specifically in ggml_gemm_i2_i8_s when called with nr > 1 (multiple activation columns).
Reproduction
# Build with AVX2 disabled (forces scalar fallback)
cmake -B build -DBITNET_ARM_TL1=OFF -DBITNET_X86_TL2=OFF
cmake --build build --target llama-server -j8
# Short prompt works
curl http://localhost:8081/v1/completions \
-d '{"model":"bitnet","prompt":"Hello","max_tokens":20}'
# → " there! I'm happy to help" ✓
# Longer prompt fails
curl http://localhost:8081/v1/completions \
-d '{"model":"bitnet","prompt":"The weather today is","max_tokens":20}'
# → "?????" ✗
Also requires fixes #588 (ReLU²) and PR #616 (weight scale direction) for coherent F16 output.
Environment
- CPU: Intel Xeon E5-2690 v2 (Ivy Bridge, AVX only, no AVX2)
- OS: Ubuntu 24.04 LTS
- Compiler: Clang 18.1.3
- CMake flags:
-DBITNET_ARM_TL1=OFF -DBITNET_X86_TL2=OFF - Model: BitNet-b1.58-2B-4T (I2_S format)
- BitNet commit: 390c30775
- 主要语言
- C++
- 星标
- 40.3k
- 派生
- 3.7k
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/BitNet 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed row 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
-
难度 2/5 1-3 小时 新手友好度 86/100
-
难度 1/5 1 小时以内 新手友好度 92/100
-
难度 2/5 1-3 小时 新手友好度 76/100
相似的 Issue
-
bug build
难度 1/5 1 小时以内 新手友好度 91/100
facebookincubator/velox#19194 ·
-
JIT-compiled number -> Decimal conversion silently overflows instead of raising DECIMAL_OVERFLOW 未关闭fuzz
难度 2/5 1-3 小时 新手友好度 82/100
ClickHouse/ClickHouse#122114 ·
-
难度 2/5 1-3 小时 新手友好度 84/100
-
module/agent platform/macos type/bug/regression
难度 2/5 1-3 小时 新手友好度 88/100
-
enhancement PyCDE
难度 2/5 1-3 小时 新手友好度 78/100