Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

vLLM oracle: GPU page fault in ROCm Triton paged-attention fallback (Qwen3.5 GDN, IQ3_XXS, gfx1200)

Đang mở
#3,167 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
32/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp, docker

Hướng nghiên cứu

Đọc .agents/oracles/vllm-gguf-plugin.md và tái hiện thiết lập ROCm/gfx1200 được ghim được mô tả ở đây. Sử dụng AMD_LOG_LEVEL/ROCM_DEBUG, sau đó nếu có thể hãy thu gọn lỗi thành một repro nhỏ hơn ở cấp Triton. Hoàn tất nghĩa là xác định được nguyên nhân có thể tái hiện ở cấp kernel hoặc một repro được tối giản; giữ gateable = no cho đến khi tổ hợp này phát ra một token.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Row: -

Found running ORACLE-VLLM-ROCM-GFX1200-DOCKER (#2961), against the pinned
vLLM oracle (e126687a9a8, ROCm/gfx1200, RX 9060 XT), loading
unsloth/Qwen3.8-27B-GGUF:UD-IQ3_XXS (a Qwen3.5-family Gated-Delta-Net
architecture) plus its mmproj-F16.gguf.

With language_model_only=True (skips multimodal dummy-data profiling) and
enforce_eager=True (skips CUDA-graph capture, whose own "minimal"
profiling KV buffer OOMs independently -- see the row's spec, ## Risks,
for that separate finding), model loading and KV-cache sizing both complete
successfully:

Available KV cache memory: 0.84 GiB
GPU KV cache size: 5,802 tokens, Maximum concurrency for 4,096 tokens per request: 1.42x
Free memory on device (15.82/15.92 GiB) on startup. ... Actual usage is 12.11 GiB
for consumed memory (weights + non-torch), 1.38 GiB for peak activation, and 0.0 GiB
for CUDAGraph memory.

Then, on the very next step:

WARNING [chunked_prefill_paged_decode.py:433] Cannot use ROCm custom paged attention
kernel, falling back to Triton implementation.
Memory access fault by GPU node-1 (Agent handle: 0x247ec020) on address 0x7f229fb44000.
Reason: Page not present or supervisor privilege.

This is a raw HIP-level trap -- no Python traceback names a line, and no
token is ever produced. This is vLLM's/Triton's own ROCm kernel dispatch,
not vllm.cpp code, so it is not fixed in the same flow here; root-causing
a page fault with no Python-level stack needs its own investigation
(AMD_LOG_LEVEL/ROCM_DEBUG, or reducing to a smaller Triton-level repro
off the full 27B model) rather than a guess landed in a hurry. Filed so
.agents/oracles/vllm-gguf-plugin.md has something concrete to point at
when it records this device: this combination has never emitted a token on
gfx1200, which is exactly the fact that keeps that file's gateable = no
correct rather than stale.

Ngôn ngữ chính
C++
Star
423
Fork
53
Merge trung bình
1 ngày 7 giờ
Pull request đã merge (30 ngày)
380

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mudler/vllm.cpp

Tất cả issue của mudler/vllm.cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.