[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 30/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- c, linux
- Lĩnh vực
- computer-graphics, operating-systems
Hướng nghiên cứu
Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.
Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.
Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.
Environment
| Kernel | 7.2.5-3-omarchy |
| amdgpu srcversion | 0321F342C8CB52128092159 |
| GPU | 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8 |
| CPU | Ryzen 7 3800X |
| Userland | torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm) |
The sysfs change
ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:
/sys/class/kfd/kfd_topology -> MISSING
This kernel exposes it at a new location instead:
/sys/class/kfd/kfd/topology/nodes/1/name -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/ -> EMPTY (0 entries)
name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.
The misdetection
$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20
gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.
I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:
baseline: name='AMD Radeon Graphics' gcnArchName='gfx1010' CUs=20
+system libdrm: name='AMD Radeon RX 5700 XT' gcnArchName='gfx1010' CUs=20
Downstream failure
The wheel's compiled arch list contains neither the reported nor the real arch:
torch.cuda.get_arch_list() ->
['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']
So the first kernel launch fails with:
RuntimeError: HIP error: invalid device function
The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:
amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0: Process python pid 3017646
amdgpu 0000:0b:00.0: in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0: Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0: PERMISSION_FAULTS: 0x3
followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:
amdgpu 0000:0b:00.0: GPU reset begin!. Source: 4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!
The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).
Questions
- Is the empty
properties/directory intended, or is property population gated on
something that is not set here? - Should
/sys/class/kfd/kfd_topology/node*/be kept as a compatibility symlink, given
that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring? - Independently: amdgpu escalating a
GCVM_L2_PROTECTION_FAULTfrom the TCP client into an
unrecoverable sdma0 hang looks like a gap inamdgpu_recover()— should that fault type
trigger a full reset sooner? (Related in spirit to #225 and #227.)
Repro
On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:
python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"
Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.
- Ngôn ngữ chính
- C
- Star
- 463
- Fork
- 145
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của ROCm/amdgpu
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
dkfans/keeperfx#5415 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
void-linux/void-runit#141 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
ARM-software/sysarch-acs#600 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày