Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Windows][ROCm/HIP] ACCESS_VIOLATION (0xc0000005) on the first decode sample when async scheduling is enabled

オープン
#3,307 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
68/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
cpp
領域
backend

調査の方向性

Start at src/vllm/v1/worker/gpu/runner.cpp:518-519 and sample_tokens_async around 5569-5684, then read SupportsAsyncSampledTokenReadback in src/vt/rocm/rocm_backend.hip and its contract in include/vt/backend.h. Compare the existing DownloadCommittedIds path and AsyncGPUModelRunnerOutput::get_output(). Done means the Windows HIP decode request no longer host-reads device memory and completes with async scheduling enabled.

索引モデルが issue の本文から書いたものです。

説明

Summary

On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.

The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.

Environment
  • Windows 11 (build 26200), Strix Halo (Radeon 8060S), gfx1151, iGPU / UMA, 128 GB shared
  • vllm.cpp commit 636926736c0b6053eda99d08c1f753b29938fac0 (also reproduced on
    9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6)
  • built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
    (C:\TheRock\10.0.0-official\build), clang/lld from TheRock, -O3 -DNDEBUG -g -Xclang -gcodeview, vllm.cpp 0.0.3 c-abi=29
  • model: Qwen3.5-4B-Q4_K_M GGUF (bartowski), --max-model-len 2048, --max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
  --host 127.0.0.1 --port 18125 --served-model-name smoke \
  --device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
  --max-num-seqs 1 --max-num-batched-tokens 512 \
  --disable-metrics --no-enable-thinking --verbose

then one request:

POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected

A chat completion with generated tokens.

Actual

Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).

Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS:  00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async

!vprot 0x6b08cb000
 State: 00002000  MEM_RESERVE
 Protect: 00000001 PAGE_NOACCESS
 Type: 00020000   MEM_PRIVATE
 RegionSize: 000000045b620000   (17.428 GB)

Faulting instruction (rdi = 0, rax = 0x6b08cb000):

140fb6180: movq 0x4f0(%rbp), %rax        ; rax = dev_ids  (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax      ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4)      ; last_sampled_tokens[i] = (int32)ids[i]

Symbolicated stack:

vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause

The device-resident decode path:

  1. src/vllm/v1/worker/gpu/runner.cpp:5569-5573 — the slot's device buffer is handed to the
    sampler, which writes the argmax ids device-resident.
  2. Both device write-back branches are inside #ifdef VLLM_CPP_CUDA
    (runner.cpp:5604-5661), so a HIP build compiles them out.
  3. The fallback else branch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
    and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids);   // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
  ...
  input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]);  // 5680-5681
}

dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.

That branch is reachable because the capability gate passes:

  • src/vllm/v1/worker/gpu/runner.cpp:112-115 — QueueSupportsAsyncInputCombine() asks
    backend->SupportsAsyncSampledTokenReadback(); runner.cpp:518-519 turns that into
    async_input_combine_.
  • src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }

while include/vt/backend.h:217-220 documents the precondition:

TODO(rocm): an INTEGRATED non-CUDA GPU reports UnifiedMemory()==true … such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.

  • the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; }               // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606

and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.

So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.

Why Linux ROCm does not see this

Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.

Workaround (verified)

VT_ASYNC_RUNNER=0 on the same binary, model and port:

HTTP_OK 'OK'  |  prefill 18/18  |  2 tokens  |  finish_reason=stop

The crash reproduces with the default (async on) 3/3 across the two commits.

Suggested fix
  1. Minimal/honest capability: SupportsAsyncSampledTokenReadback() should not be a constant
    true on ROCm — at least return unified_memory_; (same source as
    DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
    mirror/D2H exists.
  2. Implement the contract for HIP: reuse DownloadCommittedIds (runner.cpp:5234, already
    used at runner.cpp:5720 — copy queue + fork/ready events + staging buffer) or read
    slot->pinned_host after ready_event, exactly as
    AsyncGPUModelRunnerOutput::get_output() does
    (src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencing dev_ids
    on the host.
  3. Defensive: gate the host read at runner.cpp:5669 on
    backend.DeviceMemoryIsHostAddressable() and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)

Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).

Artifacts
  • full WER minidump available on request (5.5 GB, not attachable): stack above is from it
  • vllm-server.pdb (155,807,744 B) exists for the 6369267 build, so the dump symbolizes
    directly

Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).

主要言語
C++
スター
423
フォーク
53
平均マージ
1日 5時間
マージ済み PR(30日)
376

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/vllm.cpp のほかの issue

mudler/vllm.cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。