Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly

Đang mở
#968 4 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
55/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp

Hướng nghiên cứu

Start by reproducing the 344-token prompt from the issue on an RTX 5090 with MMQ enabled, then compare it with STRATA_PREFILL_MMQ=0 and an RTX 3090. Read docs/UNSLOTH_Q4.md and the context around issues #402, #420, and #954; done means identifying the sm_120 MMQ prefill fault and confirming coherent prompts without the workaround.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

On an RTX 5090, any prompt longer than a few dozen tokens comes out garbled for the model when Strata reads it through the MMQ prompt path with Unsloth's UD-IQ4_XS. The model calls the input "corrupted or fragmented", misses a planted fact, and summarizes code as the wrong kind of program. Decoding is fine: 1,500-token answers to short prompts are coherent and correct. STRATA_PREFILL_MMQ=0 fixes it at the same prompt speed, and the same binary on an RTX 3090 (sm_86) reads the same prompts correctly with MMQ on.

Setup

  • Engine 0.1.39, source build from main 6f32ec0 by setup (BUILD.json: archs [86, 120], CUDA 13.2), STRATA_MMQ_KQUANTS=OFF (setup's default)
  • RTX 5090 32 GB (sm_120) + RTX 3090 24 GB (sm_86), driver 580.119.02, Ryzen 9 9900X, 128 GB, Linux 6.18
  • --family unsloth --model UD-IQ4_XS --gguf-dir on the three shards (SHA-256s match the table in docs/UNSLOTH_Q4.md), packed by setup with --compat-bf16
  • Config as setup wrote it: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp ... --max-context 32768 --kv int8, no images

Repro

Any passage plus a question about it, temperature 0, thinking off:

txt=$(head -c 1200 docs/MULTI_GPU.md | tr '\n' ' ')
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d "$(jq -n --arg f "$txt" \
  '{model:"x",messages:[{role:"user",content:("Here is a passage:\n\n"+$f+"\n\nIn one sentence, what is this passage about? Then quote its first five words exactly.")}],max_tokens:100,temperature:0,reasoning_effort:"none"}')" \
  | jq -r .choices[0].message.content

5090, MMQ on (default), 344-token prompt:

The text provided is a corrupted or nonsensical string that appears to be a mix of: 1. Technical jargon (e.g., "GPU," "CUDA," ...

5090 with STRATA_PREFILL_MMQ=0, and the 3090 with MMQ on, same prompt:

This passage explains how Strata implements pipeline parallelism to run a single model across multiple NVIDIA GPUs by splitting layers and managing expert caches. Strata on two or three GPUs

What I measured

GPU MMQ 113 tok 344 tok 19,440-token needle (a passphrase planted mid-file)
5090 on "your message was cut off" "a corrupted or nonsensical string" missed: "incomplete or corrupted"
5090 + 3090 split (auto, K=47) on "a prompt injection attempt" "your message was cut off" missed: "corrupted or fragmented"
5090 off correct correct found
3090 on correct correct found
  • Prompts of 29-41 tokens come out right on the split with MMQ on, so the break sits somewhere between 41 and 113 tokens.
  • Prompt speed with MMQ off is the same as with it on: 690 vs 695 tok/s on the same 19K prompt on the 5090.
  • Decode with MMQ off: 121 tok/s on a 1,500-token code answer, MTP accepting 1,024 of 1,291 drafts.

The pack's experts are IQ3_S gate/up (IQ4_XS in one layer) with IQ4_NL down (Q8_0 in five layers). I only have this pack, so I can't say whether the native IQ packs hit it too. #402 ran UD-IQ4_XS on an sm_120 card at 0.1.31, but its needle check covered UD-Q4_K_XL only. #420 and #954 are other sm_120 problems in the MMQ prompt path.

Workaround: "env": {"STRATA_PREFILL_MMQ": "0"} in strata-<model>.json, or the variable in the server's environment.

Ngôn ngữ chính
C++
Star
11.6k
Fork
1k
Merge trung bình
7 giờ 46 phút
Pull request đã merge (30 ngày)
30

Chuẩn bị môi trường

Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của Niko1221/Strata

Tất cả issue của Niko1221/Strata

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.