Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU

Abierto
#233 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
30/100
Tipo de issue
Error
Claridad
Necesita aclaración
Estado de actividad
Activo
Stack tecnológico
c, linux

Línea de trabajo

Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Summary

On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.

Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.

Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.

Environment

Kernel 7.2.5-3-omarchy
amdgpu srcversion 0321F342C8CB52128092159
GPU 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8
CPU Ryzen 7 3800X
Userland torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm)

The sysfs change

ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:

/sys/class/kfd/kfd_topology                        -> MISSING

This kernel exposes it at a new location instead:

/sys/class/kfd/kfd/topology/nodes/1/name            -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id          -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/     -> EMPTY (0 entries)

name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.

The misdetection

$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
    print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20

gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.

I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:

baseline:        name='AMD Radeon Graphics'          gcnArchName='gfx1010'  CUs=20
+system libdrm:  name='AMD Radeon RX 5700 XT'        gcnArchName='gfx1010'  CUs=20

Downstream failure

The wheel's compiled arch list contains neither the reported nor the real arch:

torch.cuda.get_arch_list() ->
  ['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']

So the first kernel launch fails with:

RuntimeError: HIP error: invalid device function

The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:

amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0:  Process python pid 3017646
amdgpu 0000:0b:00.0:   in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0:          Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0:          PERMISSION_FAULTS: 0x3

followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:

amdgpu 0000:0b:00.0: GPU reset begin!. Source:  4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!

The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).

Questions

  1. Is the empty properties/ directory intended, or is property population gated on
    something that is not set here?
  2. Should /sys/class/kfd/kfd_topology/node*/ be kept as a compatibility symlink, given
    that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring?
  3. Independently: amdgpu escalating a GCVM_L2_PROTECTION_FAULT from the TCP client into an
    unrecoverable sdma0 hang looks like a gap in amdgpu_recover() — should that fault type
    trigger a full reset sooner? (Related in spirit to #225 and #227.)

Repro

On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:

python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"

Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.

Lenguaje dominante
C
Estrellas
463
Forks
145
Métricas de merge de PR
Sin PR fusionados en 30 d

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de ROCm/amdgpu

Todos los issues de ROCm/amdgpu

Issues similares

Más issues de C

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.