Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Windows][ROCm/HIP] ACCESS_VIOLATION (0xc0000005) on the first decode sample when async scheduling is enabled

Đang mở
#3,307 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp
Lĩnh vực
backend

Hướng nghiên cứu

Start at src/vllm/v1/worker/gpu/runner.cpp:518-519 and sample_tokens_async around 5569-5684, then read SupportsAsyncSampledTokenReadback in src/vt/rocm/rocm_backend.hip and its contract in include/vt/backend.h. Compare the existing DownloadCommittedIds path and AsyncGPUModelRunnerOutput::get_output(). Done means the Windows HIP decode request no longer host-reads device memory and completes with async scheduling enabled.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.

The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.

Environment
  • Windows 11 (build 26200), Strix Halo (Radeon 8060S), gfx1151, iGPU / UMA, 128 GB shared
  • vllm.cpp commit 636926736c0b6053eda99d08c1f753b29938fac0 (also reproduced on
    9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6)
  • built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
    (C:\TheRock\10.0.0-official\build), clang/lld from TheRock, -O3 -DNDEBUG -g -Xclang -gcodeview, vllm.cpp 0.0.3 c-abi=29
  • model: Qwen3.5-4B-Q4_K_M GGUF (bartowski), --max-model-len 2048, --max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
  --host 127.0.0.1 --port 18125 --served-model-name smoke \
  --device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
  --max-num-seqs 1 --max-num-batched-tokens 512 \
  --disable-metrics --no-enable-thinking --verbose

then one request:

POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected

A chat completion with generated tokens.

Actual

Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).

Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS:  00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async

!vprot 0x6b08cb000
 State: 00002000  MEM_RESERVE
 Protect: 00000001 PAGE_NOACCESS
 Type: 00020000   MEM_PRIVATE
 RegionSize: 000000045b620000   (17.428 GB)

Faulting instruction (rdi = 0, rax = 0x6b08cb000):

140fb6180: movq 0x4f0(%rbp), %rax        ; rax = dev_ids  (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax      ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4)      ; last_sampled_tokens[i] = (int32)ids[i]

Symbolicated stack:

vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause

The device-resident decode path:

  1. src/vllm/v1/worker/gpu/runner.cpp:5569-5573 — the slot's device buffer is handed to the
    sampler, which writes the argmax ids device-resident.
  2. Both device write-back branches are inside #ifdef VLLM_CPP_CUDA
    (runner.cpp:5604-5661), so a HIP build compiles them out.
  3. The fallback else branch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
    and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids);   // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
  ...
  input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]);  // 5680-5681
}

dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.

That branch is reachable because the capability gate passes:

  • src/vllm/v1/worker/gpu/runner.cpp:112-115 — QueueSupportsAsyncInputCombine() asks
    backend->SupportsAsyncSampledTokenReadback(); runner.cpp:518-519 turns that into
    async_input_combine_.
  • src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }

while include/vt/backend.h:217-220 documents the precondition:

TODO(rocm): an INTEGRATED non-CUDA GPU reports UnifiedMemory()==true … such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.

  • the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; }               // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606

and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.

So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.

Why Linux ROCm does not see this

Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.

Workaround (verified)

VT_ASYNC_RUNNER=0 on the same binary, model and port:

HTTP_OK 'OK'  |  prefill 18/18  |  2 tokens  |  finish_reason=stop

The crash reproduces with the default (async on) 3/3 across the two commits.

Suggested fix
  1. Minimal/honest capability: SupportsAsyncSampledTokenReadback() should not be a constant
    true on ROCm — at least return unified_memory_; (same source as
    DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
    mirror/D2H exists.
  2. Implement the contract for HIP: reuse DownloadCommittedIds (runner.cpp:5234, already
    used at runner.cpp:5720 — copy queue + fork/ready events + staging buffer) or read
    slot->pinned_host after ready_event, exactly as
    AsyncGPUModelRunnerOutput::get_output() does
    (src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencing dev_ids
    on the host.
  3. Defensive: gate the host read at runner.cpp:5669 on
    backend.DeviceMemoryIsHostAddressable() and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)

Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).

Artifacts
  • full WER minidump available on request (5.5 GB, not attachable): stack above is from it
  • vllm-server.pdb (155,807,744 B) exists for the 6369267 build, so the dump symbolizes
    directly

Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).

Ngôn ngữ chính
C++
Star
423
Fork
53
Merge trung bình
1 ngày 5 giờ
Pull request đã merge (30 ngày)
376

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mudler/vllm.cpp

Tất cả issue của mudler/vllm.cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.