[Windows][ROCm/HIP] ACCESS_VIOLATION (0xc0000005) on the first decode sample when async scheduling is enabled
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
Start at src/vllm/v1/worker/gpu/runner.cpp:518-519 and sample_tokens_async around 5569-5684, then read SupportsAsyncSampledTokenReadback in src/vt/rocm/rocm_backend.hip and its contract in include/vt/backend.h. Compare the existing DownloadCommittedIds path and AsyncGPUModelRunnerOutput::get_output(). Done means the Windows HIP decode request no longer host-reads device memory and completes with async scheduling enabled.
索引モデルが issue の本文から書いたものです。
説明
Summary
On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.
The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.
Environment
- Windows 11 (build 26200), Strix Halo (Radeon 8060S),
gfx1151, iGPU / UMA, 128 GB shared - vllm.cpp commit
636926736c0b6053eda99d08c1f753b29938fac0(also reproduced on
9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6) - built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
(C:\TheRock\10.0.0-official\build), clang/lld from TheRock,-O3 -DNDEBUG -g -Xclang -gcodeview,vllm.cpp 0.0.3 c-abi=29 - model:
Qwen3.5-4B-Q4_K_MGGUF (bartowski),--max-model-len 2048,--max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
--host 127.0.0.1 --port 18125 --served-model-name smoke \
--device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
--max-num-seqs 1 --max-num-batched-tokens 512 \
--disable-metrics --no-enable-thinking --verbose
then one request:
POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected
A chat completion with generated tokens.
Actual
Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).
Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS: 00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async
!vprot 0x6b08cb000
State: 00002000 MEM_RESERVE
Protect: 00000001 PAGE_NOACCESS
Type: 00020000 MEM_PRIVATE
RegionSize: 000000045b620000 (17.428 GB)
Faulting instruction (rdi = 0, rax = 0x6b08cb000):
140fb6180: movq 0x4f0(%rbp), %rax ; rax = dev_ids (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4) ; last_sampled_tokens[i] = (int32)ids[i]
Symbolicated stack:
vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause
The device-resident decode path:
src/vllm/v1/worker/gpu/runner.cpp:5569-5573— the slot's device buffer is handed to the
sampler, which writes the argmax ids device-resident.- Both device write-back branches are inside
#ifdef VLLM_CPP_CUDA
(runner.cpp:5604-5661), so a HIP build compiles them out. - The fallback
elsebranch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids); // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
...
input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]); // 5680-5681
}
dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.
That branch is reachable because the capability gate passes:
src/vllm/v1/worker/gpu/runner.cpp:112-115—QueueSupportsAsyncInputCombine()asks
backend->SupportsAsyncSampledTokenReadback();runner.cpp:518-519turns that into
async_input_combine_.src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }
while include/vt/backend.h:217-220 documents the precondition:
TODO(rocm): an INTEGRATED non-CUDA GPU reports
UnifiedMemory()==true… such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.
- the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; } // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606
and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.
So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.
Why Linux ROCm does not see this
Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.
Workaround (verified)
VT_ASYNC_RUNNER=0 on the same binary, model and port:
HTTP_OK 'OK' | prefill 18/18 | 2 tokens | finish_reason=stop
The crash reproduces with the default (async on) 3/3 across the two commits.
Suggested fix
- Minimal/honest capability:
SupportsAsyncSampledTokenReadback()should not be a constant
trueon ROCm — at leastreturn unified_memory_;(same source as
DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
mirror/D2H exists. - Implement the contract for HIP: reuse
DownloadCommittedIds(runner.cpp:5234, already
used atrunner.cpp:5720— copy queue + fork/ready events + staging buffer) or read
slot->pinned_hostafterready_event, exactly as
AsyncGPUModelRunnerOutput::get_output()does
(src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencingdev_ids
on the host. - Defensive: gate the host read at
runner.cpp:5669on
backend.DeviceMemoryIsHostAddressable()and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)
Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).
Artifacts
- full WER minidump available on request (5.5 GB, not attachable): stack above is from it
vllm-server.pdb(155,807,744 B) exists for the6369267build, so the dump symbolizes
directly
Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).
- 主要言語
- C++
- スター
- 423
- フォーク
- 53
- 平均マージ
- 1日 5時間
- マージ済み PR(30日)
- 376
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mudler/vllm.cpp のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
mudler/vllm.cpp の issue をすべて見る
似ている issue
-
bug
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
isl-org/Open3D#7585 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
Unconfirmed bug
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
luanti-org/luanti#17605 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
area: config area: firmware priority: P2 - medium size: S type: bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
Mizithra/ActiveTerrain#16 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
grumpycoders/pcsx-redux#2171 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 2 日以内に返信