[BUG] GIM unload followed by amdgpu probe can panic in amdgpu_in_reset on MI355X
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- c, linux
- Domain
- computer-graphics, operating-systems
Research direction
Start with amdgpu_vcn_ring_end_use(), amdgpu_vcn_idle_work_handler(), amdgpu_vcn_sw_fini(), and the teardown flow in amdgpu_device.c. Trace the failed VCN initialization in vcn_v5_0_1.c and the reset guard in amdgpu_dpm_switch_power_profile(). Done means failed-probe teardown cannot leave a VCN worker using destroyed reset state, with regression coverage for this sequence.
Written by the indexing model from the issue text.
Description
Summary
On an 8x AMD Instinct MI355X host, unloading GIM and then loading amdgpu can
leave a VCN idle worker running after a failed GPU probe. The host then panics
in amdgpu_in_reset() while the worker performs VCN power-profile cleanup.
The transition is reproducible with no VMs running and all VFs unbound from
vfio-pci:
- Boot with GIM loaded.
- Unbind the VFs from
vfio-pci. - Run
modprobe -r gim. - Run
modprobe amdgpu.
The first GPU probe reports a VCN ring-test timeout (-110). A later probe
then panics in amdgpu_in_reset() from amdgpu_vcn_idle_work_handler.
GIM performs a whole-GPU reset during its teardown on this platform, so that
reset may be the reason for the initial VCN timeout. However, the NULL
dereference is in the amdgpu failed-probe/cleanup path and should not be
possible even when hardware initialization fails.
Reproduction
The following was performed after all VFs were unbound and with no guest or
workload using the devices:
sudo modprobe -r gim
sudo modprobe amdgpu
The modprobe amdgpu command does not return normally; the host becomes
unresponsive and reboots. Before this, writing 0 to sriov_numvfs was also
attempted, but GIM correctly rejects SR-IOV configuration through sysfs. That
is not required for the panic reproduction.
Trace
The first PF failed VCN initialization:
amdgpu 0000:dc:00.0: [drm:amdgpu_ring_test_helper] *ERROR* ring vcn_unified_0 test failed (-110)
amdgpu 0000:dc:00.0: hw_init of IP block <vcn_v5_0_1> failed -110
amdgpu 0000:dc:00.0: amdgpu_device_ip_init failed
amdgpu 0000:dc:00.0: Fatal error during GPU init
amdgpu 0000:dc:00.0: finishing device.
amdgpu 0000:dc:00.0: probe with driver amdgpu failed with error -110
During the next PF probe, the VCN idle worker dereferenced a NULL reset
domain:
BUG: kernel NULL pointer dereference, address: 0000000000000038
Workqueue: events amdgpu_vcn_idle_work_handler [amdgpu]
RIP: 0010:amdgpu_in_reset+0x10/0x20 [amdgpu]
CR2: 0000000000000038
...
amdgpu_dpm_switch_power_profile+0x38/0xb0 [amdgpu]
amdgpu_vcn_put_profile+0x6f/0xc0 [amdgpu]
...
Unloaded tainted modules: gim(OE):2 [last unloaded: gim(OE)]
The disassembly at amdgpu_in_reset() loads adev->reset_domain and then
dereferences offset 0x38; the register holding reset_domain is zero. The
full vmcore and vmcore-dmesg.txt were retained on the test host, but the
binary vmcore is not attached to this issue.
Source analysis
The relevant control flow in the amdgpu source is:
vcn_v5_0_1_hw_init()callsamdgpu_ring_test_helper()and returns-110
before setting the IP block'sstatus.hwflag.amdgpu_device_ip_fini_early()only calls an IP block'shw_fini()when
status.hwis set. Thereforevcn_v5_0_1_hw_fini()is skipped after this
failed initialization.- The VCN idle work is scheduled by
amdgpu_vcn_ring_end_use(). amdgpu_vcn_sw_fini()does not cancel the idle work.- Later software teardown frees
adev->reset_domain. - The queued
amdgpu_vcn_idle_work_handler()calls
amdgpu_vcn_put_profile(), which calls the DPM power-profile path.
The installed Fedora amdgpu module contains an additional reset guard in
amdgpu_dpm_switch_power_profile() which calls amdgpu_in_reset(). The local
amdgpu source tree reviewed for this report does not contain that guard, but
the installed module reaches it at the reported offset. That guard also
assumes that adev->reset_domain still exists, which is not true after the
failed-probe teardown.
Relevant source files/functions:
drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c
vcn_v5_0_1_hw_init(), vcn_v5_0_1_hw_fini()
drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
amdgpu_device_ip_hw_init_phase2(), amdgpu_device_ip_fini_early(),
amdgpu_device_fini_sw()
drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
amdgpu_vcn_ring_end_use(), amdgpu_vcn_idle_work_handler(),
amdgpu_vcn_put_profile(), amdgpu_vcn_sw_fini()
drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c
ring commit/undo paths that call amdgpu_vcn_ring_end_use()
drivers/gpu/drm/amd/pm/amdgpu_dpm.c
amdgpu_dpm_switch_power_profile()
Possible fixes
Please consider:
- cancelling VCN idle work on all failed-probe/software-fini paths, rather
than only from the normalhw_finipath; - making
amdgpu_vcn_put_profile()return without entering DPM when no
workload profile is active; - making the reset guard NULL-safe before calling
amdgpu_in_reset(); and - adding regression coverage for VCN
hw_initfailure followed by teardown,
verifying that no VCN worker executes after reset-domain destruction.
The initial VCN -110 may be a separate GIM/amdgpu handoff or reset issue.
Regardless of its cause, the failed probe must not result in a worker using
freed or NULL teardown state.
Environment
| Component | Version / value |
|---|---|
| System | Dell XE9785L, 8x AMD Instinct MI355X OAM |
| OS | Fedora Linux 44 (Server Edition) |
| Kernel | 7.2.7-200.fc44.x86_64 (#1 SMP PREEMPT_DYNAMIC, built 2026-09-21) |
| CPU architecture | x86_64 |
| GPU PCI ID | 1002:75a3, revision 0 |
| GIM RPM | gim-dkms-9.2.0.K-0.noarch |
| GIM module | 9.2.0.K, srcversion 67F5729B75FC431C5E9DDA3, vermagic 7.2.7-200.fc44.x86_64 SMP preempt mod_unload |
| GIM module path | /lib/modules/7.2.7-200.fc44.x86_64/extra/gim.ko.xz |
| AMD SMI RPM | amdsmi-7.1.1-5.fc44.x86_64 |
| AMD SMI tool | amd-smi 37.0.5 |
| AMD SMI library | 64.0.1 |
| AMD SMI driver report | 9.2.0.K |
| AMD SMI path | /usr/local/bin/amd-smi |
| GPU firmware RPM | amd-gpu-firmware-20260916-1.fc44.noarch |
| Linux firmware RPM | linux-firmware-20260916-1.fc44.noarch |
| AMD microcode RPM | amd-ucode-firmware-20260916-1.fc44.noarch |
| IFWI name | AMD MI355X_36G |
| IFWI build date | 2026/07/08 01:49 |
| IFWI version field | N/A |
| GPU part number | 113-M355-01-1K1-040C |
| VBIOS | 023.040.001.008.000001, build 00193069, dated 2026/07/08 |
| VCN firmware | ENC 1.12, DEC 9, VEP 0, revision 26 |
| SDMA firmware | Version 14 (SDMA0 through SDMA3 in GIM logs) |
| Accelerator partition | DPX |
| Memory partition | NPS2 |
| VFs | 1 supported/reported and 1 enabled per GPU under this GIM configuration |
| Board model | 102-G36216-0C / 102-G36217-0C (board-dependent across the host) |
| Product | Instinct MI355 OAM |
| amdgpu module | Fedora kernel module from /lib/modules/7.2.7-200.fc44.x86_64; modinfo did not report a separate module version/srcversion |
| amdgpu source reviewed | repository amdgpu, branch master, commit 820212794d184710298efcecce42c0215bfc9c6a (amdgpu-31.60) |
| GIM source reviewed | repository MxGPU-Virtualization, branch staging, commit 7ebc33d, release 9.2.0.K |
All eight MI355X devices reported the same GPU IFWI/VBIOS, VCN firmware, DPX,
and NPS2 values unless noted above.
Questions
- Is live GIM-to-amdgpu handoff supported on MI355X, or should the drivers
reject/avoid this transition explicitly? - Is the VCN timeout after the GIM reset a known issue tracked separately?
- If this cleanup path is already fixed in another branch, which commit
should be backported to the Fedora/kernel branch containing the reset guard?
- Dominant language
- C
- Stars
- 463
- Forks
- 145
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from ROCm/amdgpu
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
Similar issues
-
rc_runtime_activate_richpresence leaves a half-initialised entry when the buffer allocation failsOpen
Difficulty 1/5 Under an hour Newbie friendliness 88/100
RetroAchievements/rcheevos#558 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
libsdl-org/SDL#16464 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
chore(gateway): emit INFO budget reserved/settled logs for proactivity v2 (chip task_2855f4ec)Possibly taken A pull request linked to this issue is open or already merged. Openbackend
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
BasedHardware/omi#20940 ·
Maintainers usually reply within 1 day