Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU

Open
#233 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
30/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Active
Tech stack
c, linux

Research direction

Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.

Written by the indexing model from the issue text.

Description

Summary

On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.

Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.

Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.

Environment

Kernel 7.2.5-3-omarchy
amdgpu srcversion 0321F342C8CB52128092159
GPU 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8
CPU Ryzen 7 3800X
Userland torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm)

The sysfs change

ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:

/sys/class/kfd/kfd_topology                        -> MISSING

This kernel exposes it at a new location instead:

/sys/class/kfd/kfd/topology/nodes/1/name            -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id          -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/     -> EMPTY (0 entries)

name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.

The misdetection

$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
    print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20

gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.

I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:

baseline:        name='AMD Radeon Graphics'          gcnArchName='gfx1010'  CUs=20
+system libdrm:  name='AMD Radeon RX 5700 XT'        gcnArchName='gfx1010'  CUs=20

Downstream failure

The wheel's compiled arch list contains neither the reported nor the real arch:

torch.cuda.get_arch_list() ->
  ['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']

So the first kernel launch fails with:

RuntimeError: HIP error: invalid device function

The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:

amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0:  Process python pid 3017646
amdgpu 0000:0b:00.0:   in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0:          Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0:          PERMISSION_FAULTS: 0x3

followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:

amdgpu 0000:0b:00.0: GPU reset begin!. Source:  4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!

The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).

Questions

  1. Is the empty properties/ directory intended, or is property population gated on
    something that is not set here?
  2. Should /sys/class/kfd/kfd_topology/node*/ be kept as a compatibility symlink, given
    that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring?
  3. Independently: amdgpu escalating a GCVM_L2_PROTECTION_FAULT from the TCP client into an
    unrecoverable sdma0 hang looks like a gap in amdgpu_recover() — should that fault type
    trigger a full reset sooner? (Related in spirit to #225 and #227.)

Repro

On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:

python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"

Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.

Dominant language
C
Stars
463
Forks
145
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from ROCm/amdgpu

All issues in ROCm/amdgpu

Similar issues

More C issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.