Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Windows][ROCm/HIP] ACCESS_VIOLATION (0xc0000005) on the first decode sample when async scheduling is enabled

Abierto
#3,307 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
68/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
cpp
Área
backend

Línea de trabajo

Start at src/vllm/v1/worker/gpu/runner.cpp:518-519 and sample_tokens_async around 5569-5684, then read SupportsAsyncSampledTokenReadback in src/vt/rocm/rocm_backend.hip and its contract in include/vt/backend.h. Compare the existing DownloadCommittedIds path and AsyncGPUModelRunnerOutput::get_output(). Done means the Windows HIP decode request no longer host-reads device memory and completes with async scheduling enabled.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Summary

On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.

The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.

Environment
  • Windows 11 (build 26200), Strix Halo (Radeon 8060S), gfx1151, iGPU / UMA, 128 GB shared
  • vllm.cpp commit 636926736c0b6053eda99d08c1f753b29938fac0 (also reproduced on
    9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6)
  • built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
    (C:\TheRock\10.0.0-official\build), clang/lld from TheRock, -O3 -DNDEBUG -g -Xclang -gcodeview, vllm.cpp 0.0.3 c-abi=29
  • model: Qwen3.5-4B-Q4_K_M GGUF (bartowski), --max-model-len 2048, --max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
  --host 127.0.0.1 --port 18125 --served-model-name smoke \
  --device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
  --max-num-seqs 1 --max-num-batched-tokens 512 \
  --disable-metrics --no-enable-thinking --verbose

then one request:

POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected

A chat completion with generated tokens.

Actual

Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).

Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS:  00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async

!vprot 0x6b08cb000
 State: 00002000  MEM_RESERVE
 Protect: 00000001 PAGE_NOACCESS
 Type: 00020000   MEM_PRIVATE
 RegionSize: 000000045b620000   (17.428 GB)

Faulting instruction (rdi = 0, rax = 0x6b08cb000):

140fb6180: movq 0x4f0(%rbp), %rax        ; rax = dev_ids  (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax      ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4)      ; last_sampled_tokens[i] = (int32)ids[i]

Symbolicated stack:

vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause

The device-resident decode path:

  1. src/vllm/v1/worker/gpu/runner.cpp:5569-5573 — the slot's device buffer is handed to the
    sampler, which writes the argmax ids device-resident.
  2. Both device write-back branches are inside #ifdef VLLM_CPP_CUDA
    (runner.cpp:5604-5661), so a HIP build compiles them out.
  3. The fallback else branch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
    and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids);   // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
  ...
  input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]);  // 5680-5681
}

dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.

That branch is reachable because the capability gate passes:

  • src/vllm/v1/worker/gpu/runner.cpp:112-115 — QueueSupportsAsyncInputCombine() asks
    backend->SupportsAsyncSampledTokenReadback(); runner.cpp:518-519 turns that into
    async_input_combine_.
  • src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }

while include/vt/backend.h:217-220 documents the precondition:

TODO(rocm): an INTEGRATED non-CUDA GPU reports UnifiedMemory()==true … such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.

  • the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; }               // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606

and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.

So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.

Why Linux ROCm does not see this

Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.

Workaround (verified)

VT_ASYNC_RUNNER=0 on the same binary, model and port:

HTTP_OK 'OK'  |  prefill 18/18  |  2 tokens  |  finish_reason=stop

The crash reproduces with the default (async on) 3/3 across the two commits.

Suggested fix
  1. Minimal/honest capability: SupportsAsyncSampledTokenReadback() should not be a constant
    true on ROCm — at least return unified_memory_; (same source as
    DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
    mirror/D2H exists.
  2. Implement the contract for HIP: reuse DownloadCommittedIds (runner.cpp:5234, already
    used at runner.cpp:5720 — copy queue + fork/ready events + staging buffer) or read
    slot->pinned_host after ready_event, exactly as
    AsyncGPUModelRunnerOutput::get_output() does
    (src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencing dev_ids
    on the host.
  3. Defensive: gate the host read at runner.cpp:5669 on
    backend.DeviceMemoryIsHostAddressable() and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)

Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).

Artifacts
  • full WER minidump available on request (5.5 GB, not attachable): stack above is from it
  • vllm-server.pdb (155,807,744 B) exists for the 6369267 build, so the dump symbolizes
    directly

Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).

Lenguaje dominante
C++
Estrellas
423
Forks
53
Merge medio
1 d 7 h
PR fusionados (30 d)
380

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de mudler/vllm.cpp

Todos los issues de mudler/vllm.cpp

Issues similares

Más issues de C++

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.