Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[BUG] KFD topology sysfs moved + properties/ empty on 7.2 -> ROCm silently misdetects Navi 10 as gfx1010/20CU; HSA_OVERRIDE_GFX_VERSION then hangs the GPU

オープン
#233 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
30/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
活発
技術スタック
c, linux

調査の方向性

Start by reproducing the Python torch command and comparing /sys/class/kfd/kfd_topology with /sys/class/kfd/kfd/topology/nodes/*/name and properties/. Read the reported KFD topology and amdgpu recovery behavior, then determine whether the issue belongs in this fork or the linked upstream GitLab project. Done requires a decided scope and a verified fix or clear upstream handoff for the sysfs and fault-recovery concerns.

索引モデルが issue の本文から書いたものです。

説明

Summary

On kernel 7.2 the KFD topology sysfs has moved and its per-node properties/ directory is
empty. Older ROCm userland (6.2) looks for the old path, finds nothing, and silently falls
back to a generic agent record — reporting a Navi 10 as gfx1010 with 20 CUs instead of
gfx103x with 40. No error is surfaced to the application.

Forcing a code object to load on top of that misdetection (HSA_OVERRIDE_GFX_VERSION=10.3.0)
produces a GCVM_L2_PROTECTION_FAULT and an unrecoverable sdma0 hang requiring a full GPU
reset.

Note: this repo's README points bug reports at https://gitlab.freedesktop.org/drm/amd/-/issues .
I did not have an account there, so filing here; the kernel-side half of this (empty
properties/, or restoring the old path as a compat symlink) probably belongs there.

Environment

Kernel 7.2.5-3-omarchy
amdgpu srcversion 0321F342C8CB52128092159
GPU 1002:731f rev 0xc1 — Navi 10 / RX 5700 XT, 8176 MiB VRAM, VBIOS 113-230LNAVIXT612_8GD6_MS_W8
CPU Ryzen 7 3800X
Userland torch 2.5.1+rocm6.2, HIP 6.2.41133-dd7f95766 (pip wheel, no /opt/rocm)

The sysfs change

ROCm 6.2-era userland reads the long-standing KFD topology path. That path is gone:

/sys/class/kfd/kfd_topology                        -> MISSING

This kernel exposes it at a new location instead:

/sys/class/kfd/kfd/topology/nodes/1/name            -> navi10
/sys/class/kfd/kfd/topology/nodes/1/gpu_id          -> 52383
/sys/class/kfd/kfd/topology/nodes/1/properties/     -> EMPTY (0 entries)

name and gpu_id are populated, but properties/ is empty on every node (both the CPU
node 0 and the GPU node 1). That directory is where isa, simd_count, cu_count, etc.
would live.

The misdetection

$ python -c "import torch; p=torch.cuda.get_device_properties(0); \
    print(p.name, p.gcnArchName, p.multi_processor_count)"
AMD Radeon Graphics gfx1010:xnack- 20

gfx1010 is Vega 20 / Radeon VII. This is Navi 10, which should report gfx103x with 40
CUs. Memory size (8.6 GB) is correct, because it is read through a path that still works —
so the record is partly right, which is what makes it dangerous.

I confirmed the ids-file path is a red herring: the wheel's bundled
lib/libdrm_amdgpu.so looks for /opt/amdgpu/share/libdrm/amdgpu.ids, which does not exist,
and that only explains the generic name. Preloading the system libdrm_amdgpu.so (which
reads /usr/share/libdrm/amdgpu.ids) corrects the name but leaves gcnArchName=gfx1010 and
20 CUs unchanged:

baseline:        name='AMD Radeon Graphics'          gcnArchName='gfx1010'  CUs=20
+system libdrm:  name='AMD Radeon RX 5700 XT'        gcnArchName='gfx1010'  CUs=20

Downstream failure

The wheel's compiled arch list contains neither the reported nor the real arch:

torch.cuda.get_arch_list() ->
  ['gfx900', 'gfx906', 'gfx908', 'gfx90a', 'gfx1030', 'gfx1100', 'gfx942']

So the first kernel launch fails with:

RuntimeError: HIP error: invalid device function

The obvious workaround — HSA_OVERRIDE_GFX_VERSION=10.3.0, to make the device claim to be
gfx1030 so a code object will load — is what actually kills the GPU. Two seconds after
setting it:

amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:4798)
amdgpu 0000:0b:00.0:  Process python pid 3017646
amdgpu 0000:0b:00.0:   in page starting at address 0x000000000440e000 from client 0x1b (UTCL2)
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00801031
amdgpu 0000:0b:00.0:          Faulty UTCL2 client ID: TCP (0x8)
amdgpu 0000:0b:00.0:          PERMISSION_FAULTS: 0x3

followed by ~30 ring sdma0 timeouts over 90s and escalation to a full reset:

amdgpu 0000:0b:00.0: GPU reset begin!. Source:  4
amdgpu 0000:0b:00.0: VRAM is lost due to GPU reset!
amdgpu 0000:0b:00.0: GPU reset(36) succeeded!

The reset dropped the display server along with it (Hyprland died in
Aquamarine::CDRMBackend::flushAsyncCommitEvents during DRM backend teardown).

Questions

  1. Is the empty properties/ directory intended, or is property population gated on
    something that is not set here?
  2. Should /sys/class/kfd/kfd_topology/node*/ be kept as a compatibility symlink, given
    that in-tree-adjacent userland (ROCm) reads it and fails silently rather than erroring?
  3. Independently: amdgpu escalating a GCVM_L2_PROTECTION_FAULT from the TCP client into an
    unrecoverable sdma0 hang looks like a gap in amdgpu_recover() — should that fault type
    trigger a full reset sooner? (Related in spirit to #225 and #227.)

Repro

On any kernel with the new KFD sysfs layout plus ROCm userland old enough to expect the old
path:

python -c "import torch; print(torch.cuda.get_device_properties(0).gcnArchName)"

Compare against the name in /sys/class/kfd/kfd/topology/nodes/*/name and the PCI ID. A
gfx10xx answer that disagrees with a Navi 10 node name is the tell.

主要言語
C
スター
463
フォーク
145
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

ROCm/amdgpu のほかの issue

ROCm/amdgpu の issue をすべて見る

似ている issue

C の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。