Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[BUG] GIM unload followed by amdgpu probe can panic in amdgpu_in_reset on MI355X

Open
#234 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
c, linux

Research direction

Start with amdgpu_vcn_ring_end_use(), amdgpu_vcn_idle_work_handler(), amdgpu_vcn_sw_fini(), and the teardown flow in amdgpu_device.c. Trace the failed VCN initialization in vcn_v5_0_1.c and the reset guard in amdgpu_dpm_switch_power_profile(). Done means failed-probe teardown cannot leave a VCN worker using destroyed reset state, with regression coverage for this sequence.

Written by the indexing model from the issue text.

Description

Summary

On an 8x AMD Instinct MI355X host, unloading GIM and then loading amdgpu can
leave a VCN idle worker running after a failed GPU probe. The host then panics
in amdgpu_in_reset() while the worker performs VCN power-profile cleanup.

The transition is reproducible with no VMs running and all VFs unbound from
vfio-pci:

  1. Boot with GIM loaded.
  2. Unbind the VFs from vfio-pci.
  3. Run modprobe -r gim.
  4. Run modprobe amdgpu.

The first GPU probe reports a VCN ring-test timeout (-110). A later probe
then panics in amdgpu_in_reset() from amdgpu_vcn_idle_work_handler.

GIM performs a whole-GPU reset during its teardown on this platform, so that
reset may be the reason for the initial VCN timeout. However, the NULL
dereference is in the amdgpu failed-probe/cleanup path and should not be
possible even when hardware initialization fails.

Reproduction

The following was performed after all VFs were unbound and with no guest or
workload using the devices:

sudo modprobe -r gim
sudo modprobe amdgpu

The modprobe amdgpu command does not return normally; the host becomes
unresponsive and reboots. Before this, writing 0 to sriov_numvfs was also
attempted, but GIM correctly rejects SR-IOV configuration through sysfs. That
is not required for the panic reproduction.

Trace

The first PF failed VCN initialization:

amdgpu 0000:dc:00.0: [drm:amdgpu_ring_test_helper] *ERROR* ring vcn_unified_0 test failed (-110)
amdgpu 0000:dc:00.0: hw_init of IP block <vcn_v5_0_1> failed -110
amdgpu 0000:dc:00.0: amdgpu_device_ip_init failed
amdgpu 0000:dc:00.0: Fatal error during GPU init
amdgpu 0000:dc:00.0: finishing device.
amdgpu 0000:dc:00.0: probe with driver amdgpu failed with error -110

During the next PF probe, the VCN idle worker dereferenced a NULL reset
domain:

BUG: kernel NULL pointer dereference, address: 0000000000000038
Workqueue: events amdgpu_vcn_idle_work_handler [amdgpu]
RIP: 0010:amdgpu_in_reset+0x10/0x20 [amdgpu]
CR2: 0000000000000038
...
amdgpu_dpm_switch_power_profile+0x38/0xb0 [amdgpu]
amdgpu_vcn_put_profile+0x6f/0xc0 [amdgpu]
...
Unloaded tainted modules: gim(OE):2 [last unloaded: gim(OE)]

The disassembly at amdgpu_in_reset() loads adev->reset_domain and then
dereferences offset 0x38; the register holding reset_domain is zero. The
full vmcore and vmcore-dmesg.txt were retained on the test host, but the
binary vmcore is not attached to this issue.

Source analysis

The relevant control flow in the amdgpu source is:

  • vcn_v5_0_1_hw_init() calls amdgpu_ring_test_helper() and returns -110
    before setting the IP block's status.hw flag.
  • amdgpu_device_ip_fini_early() only calls an IP block's hw_fini() when
    status.hw is set. Therefore vcn_v5_0_1_hw_fini() is skipped after this
    failed initialization.
  • The VCN idle work is scheduled by amdgpu_vcn_ring_end_use().
  • amdgpu_vcn_sw_fini() does not cancel the idle work.
  • Later software teardown frees adev->reset_domain.
  • The queued amdgpu_vcn_idle_work_handler() calls
    amdgpu_vcn_put_profile(), which calls the DPM power-profile path.

The installed Fedora amdgpu module contains an additional reset guard in
amdgpu_dpm_switch_power_profile() which calls amdgpu_in_reset(). The local
amdgpu source tree reviewed for this report does not contain that guard, but
the installed module reaches it at the reported offset. That guard also
assumes that adev->reset_domain still exists, which is not true after the
failed-probe teardown.

Relevant source files/functions:

drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c
  vcn_v5_0_1_hw_init(), vcn_v5_0_1_hw_fini()
drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
  amdgpu_device_ip_hw_init_phase2(), amdgpu_device_ip_fini_early(),
  amdgpu_device_fini_sw()
drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
  amdgpu_vcn_ring_end_use(), amdgpu_vcn_idle_work_handler(),
  amdgpu_vcn_put_profile(), amdgpu_vcn_sw_fini()
drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c
  ring commit/undo paths that call amdgpu_vcn_ring_end_use()
drivers/gpu/drm/amd/pm/amdgpu_dpm.c
  amdgpu_dpm_switch_power_profile()

Possible fixes

Please consider:

  • cancelling VCN idle work on all failed-probe/software-fini paths, rather
    than only from the normal hw_fini path;
  • making amdgpu_vcn_put_profile() return without entering DPM when no
    workload profile is active;
  • making the reset guard NULL-safe before calling amdgpu_in_reset(); and
  • adding regression coverage for VCN hw_init failure followed by teardown,
    verifying that no VCN worker executes after reset-domain destruction.

The initial VCN -110 may be a separate GIM/amdgpu handoff or reset issue.
Regardless of its cause, the failed probe must not result in a worker using
freed or NULL teardown state.

Environment

Component Version / value
System Dell XE9785L, 8x AMD Instinct MI355X OAM
OS Fedora Linux 44 (Server Edition)
Kernel 7.2.7-200.fc44.x86_64 (#1 SMP PREEMPT_DYNAMIC, built 2026-09-21)
CPU architecture x86_64
GPU PCI ID 1002:75a3, revision 0
GIM RPM gim-dkms-9.2.0.K-0.noarch
GIM module 9.2.0.K, srcversion 67F5729B75FC431C5E9DDA3, vermagic 7.2.7-200.fc44.x86_64 SMP preempt mod_unload
GIM module path /lib/modules/7.2.7-200.fc44.x86_64/extra/gim.ko.xz
AMD SMI RPM amdsmi-7.1.1-5.fc44.x86_64
AMD SMI tool amd-smi 37.0.5
AMD SMI library 64.0.1
AMD SMI driver report 9.2.0.K
AMD SMI path /usr/local/bin/amd-smi
GPU firmware RPM amd-gpu-firmware-20260916-1.fc44.noarch
Linux firmware RPM linux-firmware-20260916-1.fc44.noarch
AMD microcode RPM amd-ucode-firmware-20260916-1.fc44.noarch
IFWI name AMD MI355X_36G
IFWI build date 2026/07/08 01:49
IFWI version field N/A
GPU part number 113-M355-01-1K1-040C
VBIOS 023.040.001.008.000001, build 00193069, dated 2026/07/08
VCN firmware ENC 1.12, DEC 9, VEP 0, revision 26
SDMA firmware Version 14 (SDMA0 through SDMA3 in GIM logs)
Accelerator partition DPX
Memory partition NPS2
VFs 1 supported/reported and 1 enabled per GPU under this GIM configuration
Board model 102-G36216-0C / 102-G36217-0C (board-dependent across the host)
Product Instinct MI355 OAM
amdgpu module Fedora kernel module from /lib/modules/7.2.7-200.fc44.x86_64; modinfo did not report a separate module version/srcversion
amdgpu source reviewed repository amdgpu, branch master, commit 820212794d184710298efcecce42c0215bfc9c6a (amdgpu-31.60)
GIM source reviewed repository MxGPU-Virtualization, branch staging, commit 7ebc33d, release 9.2.0.K

All eight MI355X devices reported the same GPU IFWI/VBIOS, VCN firmware, DPX,
and NPS2 values unless noted above.

Questions

  1. Is live GIM-to-amdgpu handoff supported on MI355X, or should the drivers
    reject/avoid this transition explicitly?
  2. Is the VCN timeout after the GIM reset a known issue tracked separately?
  3. If this cleanup path is already fixed in another branch, which commit
    should be backported to the Fedora/kernel branch containing the reset guard?
Dominant language
C
Stars
463
Forks
145
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from ROCm/amdgpu

All issues in ROCm/amdgpu

Similar issues

More C issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.