[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- c, linux
- Domain
- computer-graphics, operating-systems
Research direction
Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.
Written by the indexing model from the issue text.
Description
Summary
On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.
Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.
Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.
Environment
| Kernel | 7.2.5-3-omarchy |
| amdgpu srcversion | 0321F342C8CB52128092159 |
| GPU | 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8 |
| CPU | Ryzen 7 3800X |
| Userland | torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm) |
The sysfs change
ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:
/sys/class/kfd/kfd_topology -> MISSING
This kernel exposes it at a new location instead:
/sys/class/kfd/kfd/topology/nodes/1/name -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/ -> EMPTY (0 entries)
name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.
The misdetection
$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20
gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.
I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:
baseline: name='AMD Radeon Graphics' gcnArchName='gfx1010' CUs=20
+system libdrm: name='AMD Radeon RX 5700 XT' gcnArchName='gfx1010' CUs=20
Downstream failure
The wheel's compiled arch list contains neither the reported nor the real arch:
torch.cuda.get_arch_list() ->
['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']
So the first kernel launch fails with:
RuntimeError: HIP error: invalid device function
The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:
amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0: Process python pid 3017646
amdgpu 0000:0b:00.0: in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0: Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0: PERMISSION_FAULTS: 0x3
followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:
amdgpu 0000:0b:00.0: GPU reset begin!. Source: 4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!
The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).
Questions
- Is the empty
properties/directory intended, or is property population gated on
something that is not set here? - Should
/sys/class/kfd/kfd_topology/node*/be kept as a compatibility symlink, given
that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring? - Independently: amdgpu escalating a
GCVM_L2_PROTECTION_FAULTfrom the TCP client into an
unrecoverable sdma0 hang looks like a gap inamdgpu_recover()— should that fault type
trigger a full reset sooner? (Related in spirit to #225 and #227.)
Repro
On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:
python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"
Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.
- Dominant language
- C
- Stars
- 463
- Forks
- 145
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from ROCm/amdgpu
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
dkfans/keeperfx#5415 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
ARM-software/sysarch-acs#600 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
EchoTools/nevr-runtime#117 · 2 comments ·
Maintainers usually reply within 1 day