Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU

Đang mở
#233 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
30/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
c, linux

Hướng nghiên cứu

Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.

Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.

Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.

Environment

Kernel 7.2.5-3-omarchy
amdgpu srcversion 0321F342C8CB52128092159
GPU 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8
CPU Ryzen 7 3800X
Userland torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm)

The sysfs change

ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:

/sys/class/kfd/kfd_topology                        -> MISSING

This kernel exposes it at a new location instead:

/sys/class/kfd/kfd/topology/nodes/1/name            -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id          -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/     -> EMPTY (0 entries)

name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.

The misdetection

$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
    print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20

gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.

I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:

baseline:        name='AMD Radeon Graphics'          gcnArchName='gfx1010'  CUs=20
+system libdrm:  name='AMD Radeon RX 5700 XT'        gcnArchName='gfx1010'  CUs=20

Downstream failure

The wheel's compiled arch list contains neither the reported nor the real arch:

torch.cuda.get_arch_list() ->
  ['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']

So the first kernel launch fails with:

RuntimeError: HIP error: invalid device function

The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:

amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0:  Process python pid 3017646
amdgpu 0000:0b:00.0:   in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0:          Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0:          PERMISSION_FAULTS: 0x3

followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:

amdgpu 0000:0b:00.0: GPU reset begin!. Source:  4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!

The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).

Questions

  1. Is the empty properties/ directory intended, or is property population gated on
    something that is not set here?
  2. Should /sys/class/kfd/kfd_topology/node*/ be kept as a compatibility symlink, given
    that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring?
  3. Independently: amdgpu escalating a GCVM_L2_PROTECTION_FAULT from the TCP client into an
    unrecoverable sdma0 hang looks like a gap in amdgpu_recover() — should that fault type
    trigger a full reset sooner? (Related in spirit to #225 and #227.)

Repro

On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:

python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"

Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.

Ngôn ngữ chính
C
Star
463
Fork
145
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của ROCm/amdgpu

Tất cả issue của ROCm/amdgpu

Issue tương tự

Thêm issue về C

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.