[Windows][ROCm/HIP] ACCESS_VIOLATION (0xc0000005) on the first decode sample when async scheduling is enabled
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
Línea de trabajo
Start at src/vllm/v1/worker/gpu/runner.cpp:518-519 and sample_tokens_async around 5569-5684, then read SupportsAsyncSampledTokenReadback in src/vt/rocm/rocm_backend.hip and its contract in include/vt/backend.h. Compare the existing DownloadCommittedIds path and AsyncGPUModelRunnerOutput::get_output(). Done means the Windows HIP decode request no longer host-reads device memory and completes with async scheduling enabled.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.
The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.
Environment
- Windows 11 (build 26200), Strix Halo (Radeon 8060S),
gfx1151, iGPU / UMA, 128 GB shared - vllm.cpp commit
636926736c0b6053eda99d08c1f753b29938fac0(also reproduced on
9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6) - built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
(C:\TheRock\10.0.0-official\build), clang/lld from TheRock,-O3 -DNDEBUG -g -Xclang -gcodeview,vllm.cpp 0.0.3 c-abi=29 - model:
Qwen3.5-4B-Q4_K_MGGUF (bartowski),--max-model-len 2048,--max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
--host 127.0.0.1 --port 18125 --served-model-name smoke \
--device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
--max-num-seqs 1 --max-num-batched-tokens 512 \
--disable-metrics --no-enable-thinking --verbose
then one request:
POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected
A chat completion with generated tokens.
Actual
Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).
Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS: 00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async
!vprot 0x6b08cb000
State: 00002000 MEM_RESERVE
Protect: 00000001 PAGE_NOACCESS
Type: 00020000 MEM_PRIVATE
RegionSize: 000000045b620000 (17.428 GB)
Faulting instruction (rdi = 0, rax = 0x6b08cb000):
140fb6180: movq 0x4f0(%rbp), %rax ; rax = dev_ids (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4) ; last_sampled_tokens[i] = (int32)ids[i]
Symbolicated stack:
vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause
The device-resident decode path:
src/vllm/v1/worker/gpu/runner.cpp:5569-5573— the slot's device buffer is handed to the
sampler, which writes the argmax ids device-resident.- Both device write-back branches are inside
#ifdef VLLM_CPP_CUDA
(runner.cpp:5604-5661), so a HIP build compiles them out. - The fallback
elsebranch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids); // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
...
input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]); // 5680-5681
}
dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.
That branch is reachable because the capability gate passes:
src/vllm/v1/worker/gpu/runner.cpp:112-115—QueueSupportsAsyncInputCombine()asks
backend->SupportsAsyncSampledTokenReadback();runner.cpp:518-519turns that into
async_input_combine_.src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }
while include/vt/backend.h:217-220 documents the precondition:
TODO(rocm): an INTEGRATED non-CUDA GPU reports
UnifiedMemory()==true… such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.
- the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; } // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606
and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.
So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.
Why Linux ROCm does not see this
Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.
Workaround (verified)
VT_ASYNC_RUNNER=0 on the same binary, model and port:
HTTP_OK 'OK' | prefill 18/18 | 2 tokens | finish_reason=stop
The crash reproduces with the default (async on) 3/3 across the two commits.
Suggested fix
- Minimal/honest capability:
SupportsAsyncSampledTokenReadback()should not be a constant
trueon ROCm — at leastreturn unified_memory_;(same source as
DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
mirror/D2H exists. - Implement the contract for HIP: reuse
DownloadCommittedIds(runner.cpp:5234, already
used atrunner.cpp:5720— copy queue + fork/ready events + staging buffer) or read
slot->pinned_hostafterready_event, exactly as
AsyncGPUModelRunnerOutput::get_output()does
(src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencingdev_ids
on the host. - Defensive: gate the host read at
runner.cpp:5669on
backend.DeviceMemoryIsHostAddressable()and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)
Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).
Artifacts
- full WER minidump available on request (5.5 GB, not attachable): stack above is from it
vllm-server.pdb(155,807,744 B) exists for the6369267build, so the dump symbolizes
directly
Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).
- Lenguaje dominante
- C++
- Estrellas
- 423
- Forks
- 53
- Merge medio
- 1 d 7 h
- PR fusionados (30 d)
- 380
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/vllm.cpp
-
[Windows] full build fails in tools/bench/conv1d_scaling_probe.cpp (POSIX-only sys/resource.h)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
Todos los issues de mudler/vllm.cpp
Issues similares
-
Unconfirmed bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
luanti-org/luanti#17605 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
area: config area: firmware priority: P2 - medium size: S type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Mizithra/ActiveTerrain#16 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
grumpycoders/pcsx-redux#2171 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
bytedance/trae-agent#524 · 1 comentario ·
Los mantenedores suelen responder en 1 día