Resume from hibernation always fails when the driver is in the initramfs: `nv_pm_notifier` does not handle `PM_RESTORE_PREPARE`
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 72/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
- 技术栈
- c, linux
调研方向
从 kernel-open/nvidia/nv.c 中的 nv_pm_notifier 开始,并将 PM_RESTORE_PREPARE 路径与 PM_HIBERNATION_PREPARE 进行比较。确认 notifier 通过允许的 suspend 路径处理 restore 事件,然后在 initramfs 中包含 NVIDIA 模块的情况下重现 hibernation,并验证恢复的会话不再在 nv_pmops_freeze 中失败。
由索引模型根据 Issue 内容生成。
描述
Summary
When the NVIDIA modules are loaded from the initramfs (the default on Arch Linux, and
therefore on Arch-derived distributions), resume from hibernation fails 100% of the time
with nv_pmops_freeze returning -EIO, provided NVreg_PreserveVideoMemoryAllocations=1 —
which Arch's own packaging sets by default.
Writing the hibernation image always succeeds. Only the restore fails, and it fails one step
after the image has been read back correctly.
This is distinct from the Blackwell early-KMS hibernation hangs discussed in distribution
trackers: this reproduces on Turing, and it is a logic gap in the PM notifier that is visible
in the source rather than a modesetting problem.
Environment
| Driver | nvidia-open-dkms 610.57.04 |
| GPU | Quadro T2000 (Turing, TU117GLM), 0000:01:00.0 |
| Machine | Lenovo ThinkPad P53 (20QN), Optimus laptop |
| Kernel | 7.1.9-arch1-2 |
| Distro | Omarchy (Arch-based) |
| Display topology | Only connected output is eDP-1 on the Intel iGPU. nvidia-drm logs Cannot find any crtc or sizes. |
Relevant module parameters, all distro defaults:
PreserveVideoMemoryAllocations: 1 # /usr/lib/modprobe.d/gsr-nvidia.conf
UseKernelSuspendNotifiers: 1 # /usr/lib/modprobe.d/nvidia-sleep.conf
Modules early-loaded into the initramfs by /etc/mkinitcpio.conf.d/nvidia.conf:
MODULES+=(nvidia nvidia_modeset nvidia_uvm nvidia_drm)
Steps to reproduce
- Arch-based system,
nvidia-open-dkms, NVIDIA modules in the initramfsMODULESarray. NVreg_PreserveVideoMemoryAllocations=1(default ifgpu-screen-recorderis installed).- A working hibernation setup — valid
resume=andresume_offset=. systemctl hibernate.- Power the machine back on.
Expected: the session is restored.
Actual: the image is found and read back successfully, then the restore is abandoned and the
machine continues into a fresh boot.
Log
Run /init as init process
nvidia: loading out-of-tree module ... <- initramfs
[drm] Initialized nvidia-drm 0.0.0 for 0000:01:00.0 on minor 1
nvidia 0000:01:00.0: [drm] Cannot find any crtc or sizes
PM: Image signature found, resuming
PM: hibernation: Read 5563848 kbytes in 3.28 seconds (1696.29 MB/s)
PM: Image successfully loaded
NVRM: GPU 0000:01:00.0: PreserveVideoMemoryAllocations module parameter is set.
System Power Management attempted without driver procfs suspend interface. ...
nvidia 0000:01:00.0: PM: pci_pm_freeze(): nv_pmops_freeze [nvidia] returns -5
nvidia 0000:01:00.0: PM: failed to quiesce async: error -5
PM: hibernation: Failed to load image, recovering.
PM: hibernation: resume failed (-5)
Note that the image was read back at 1.7 GB/s with a valid signature — the hibernation setup
itself is entirely correct. The failure is strictly after Image successfully loaded.
Analysis
Restoring a hibernation image is a two-kernel operation. The freshly booted resume kernel
loads the image, then must quiesce its own devices —
hibernation_restore() → dpm_suspend_start(PMSG_QUIESCE) → each driver's .freeze —
before jumping into the restored image. Because the driver is in the initramfs, that resume
kernel has a live, initialised NVIDIA device to freeze.
kernel-open/nvidia/nv.c — nv_pmops_freeze() calls
nvidia_suspend(dev, NV_PM_ACTION_HIBERNATE, is_procfs_suspend=NV_FALSE), and
nvidia_suspend() contains:
if (nv->preserve_vidmem_allocations &&
nv_dev_needs_vidmem_preservation(nv) &&
!is_procfs_suspend)
{
...
status = NV_ERR_NOT_SUPPORTED;
goto done;
}
All three conditions hold on the resume path:
preserve_vidmem_allocations— set, viaNVreg_PreserveVideoMemoryAllocations=1.nv_dev_needs_vidmem_preservation()(common/inc/nv.h) returns
!is_tegra_pci_igpu && !NV_IS_SOC_DISPLAY_DEVICE, true for a discrete PCI GPU.is_procfs_suspendisNV_FALSE, because this is the kernel PM callback.
NV_ERR_NOT_SUPPORTED → nv_pmops_freeze returns -EIO → the restore is abandoned.
Why the save path does not hit this. With NVreg_UseKernelSuspendNotifiers=1 the driver
registers nv_pm_notifier. On the way down the kernel fires PM_HIBERNATION_PREPARE, the
notifier runs nv_suspend_devices(), which calls nvidia_suspend(..., is_procfs_suspend=NV_TRUE)
— the permitted path — saves video memory and sets NV_FLAG_SUSPENDED. The subsequent
nv_pmops_freeze then short-circuits on that flag and returns success.
The gap. On the way back up, software_resume() fires PM_RESTORE_PREPARE, and
nv_pm_notifier's switch handles only:
case PM_SUSPEND_PREPARE:
case PM_HIBERNATION_PREPARE:
case PM_POST_SUSPEND:
case PM_POST_HIBERNATION:
default: return NOTIFY_DONE; /* PM_RESTORE_PREPARE lands here */
PM_RESTORE_PREPARE falls through to default. Nothing pre-authorises the freeze,
NV_FLAG_SUSPENDED is clear, and the refusal above fires unconditionally.
This is deterministic, not intermittent.
Suggested fix
Handle PM_RESTORE_PREPARE in nv_pm_notifier alongside PM_HIBERNATION_PREPARE, so the
resume kernel's devices are quiesced through the same permitted path the suspend side uses.
Workaround
Keep the NVIDIA modules out of the initramfs, so the resume kernel has no NVIDIA device bound
and there is no .freeze callback to refuse:
# /etc/mkinitcpio.conf.d/zz-nvidia-no-early-load.conf (must sort last)
_mods=()
for _m in "${MODULES[@]}"; do
case $_m in nvidia|nvidia_modeset|nvidia_uvm|nvidia_drm) ;; *) _mods+=("$_m") ;; esac
done
MODULES=("${_mods[@]}")
unset _mods _m
Then rebuild the initramfs. On this machine that restored hibernation completely: entry at
11:06:39, resume at 11:07:53, no PM errors, same boot ID — a genuine restore rather than a
fresh boot.
This costs nothing on an Optimus laptop whose only connected display is on the iGPU. It would
not be acceptable on a system where NVIDIA drives the panel and early KMS is wanted, which is
why it is a workaround rather than a fix.
Impact
Any configuration that early-loads the NVIDIA modules with
NVreg_PreserveVideoMemoryAllocations=1 cannot resume from hibernation. On Arch that is the
default pairing: nvidia.conf puts the modules in the initramfs, and gsr-nvidia.conf
(shipped with gpu-screen-recorder) sets the flag.
- 主要语言
- C
- 星标
- 17.4k
- 派生
- 1.9k
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/open-gpu-kernel-modules 的其他 Issue
-
PlatformRequestHandler: missing !bInit guard on bSystemParamLimitUpdate causes NV_ERR_INVALID_DATA at boot可能已有人在做 @SammyTourani 于 10 天前认领。 未关闭
难度 1/5 1 小时以内 新手友好度 88/100
NVIDIA/open-gpu-kernel-modules#1360 ·
维护者通常 1 天内回复
-
HKC G27H7Pro redundant long-HPD loop after DisplayPort link training (RTX 5060 Laptop, 610.57.04)未关闭bug
难度 2/5 1-3 小时 新手友好度 88/100
NVIDIA/open-gpu-kernel-modules#1323 ·
维护者通常 1 天内回复
-
build-problem
难度 2/5 1-3 小时 新手友好度 74/100
NVIDIA/open-gpu-kernel-modules#1155 · 1 条评论 ·
维护者通常 1 天内回复
-
build-problem
难度 2/5 1-3 小时 新手友好度 76/100
NVIDIA/open-gpu-kernel-modules#1154 · 10 条评论 ·
维护者通常 1 天内回复
-
Bug template is outdated可能已有人在做 @Rohithmatham12 于 121 天前认领。 未关闭bug
难度 1/5 1 小时以内 新手友好度 78/100
NVIDIA/open-gpu-kernel-modules#1147 ·
维护者通常 1 天内回复
查看 NVIDIA/open-gpu-kernel-modules 的全部 Issue
相似的 Issue
-
CVE-2026-18839 popt: size_t underflow in `singleOptionHelp`可能已有人在做 @pmatilai 于 34 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 70/100
rpm-software-management/popt#143 ·
-
难度 1/5 1 小时以内 新手友好度 85/100
-
难度 1/5 1 小时以内 新手友好度 85/100
libretro/mupen64plus-libretro-nx#663 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 90/100
-
难度 2/5 1-3 小时 新手友好度 65/100
dkfans/keeperfx#5415 · 1 条评论 ·
维护者通常 1 天内回复