DCGM API calls block past 45s on a GPU generating correctable ECC errors under load, and return in 0s when it is idle
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 38/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- cpp
- Domain
- observability-sre
Research direction
No source file or test is named. Start by reproducing the behavior with dcgmi health -g 0 -c and dcgmi diag while the affected GPU is under load and idle, comparing the results with nvidia-smi and the reported ECC fields. Done means establishing the cause and documenting an actionable timeout or health-signal behavior for monitoring clients.
Written by the indexing model from the issue text.
Description
Summary
On a GB200 node with one GPU generating a very high rate of correctable ECC errors, DCGM API calls block for longer than 45s while that GPU is under load, and return in 0s once the GPU is idle. The stall is reproducible against the affected GPU and disappears entirely when the workload leaves, with no configuration change.
A monitoring agent that calls DCGM on a fixed interval is therefore repeatedly wedged by a GPU that DCGM simultaneously reports as Healthy.
Environment
| DCGM | 4.5.2 |
| Driver | 580.126.20 |
| GPU | NVIDIA GB200 (device id 2941), 4 per node |
| OS / kernel | Ubuntu 24.04.4 LTS, 6.17.0-1014-nvidia-64k |
| Mode | nv-hostengine in a container, clients connect remotely on port 5555 |
The affected GPU
One GPU of the four, consistently:
| Field | Value |
|---|---|
DCGM_FI_DEV_ECC_SBE_AGG_TOTAL |
54,731,746,511 |
DCGM_FI_DEV_ECC_SBE_VOL_TOTAL |
9.6e9, rising at ~17,700/s under load |
DCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWS |
8 |
DCGM_FI_DEV_ECC_DBE_VOL_TOTAL / _AGG_TOTAL |
0 / 0 |
DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS |
0 |
DCGM_FI_DEV_ROW_REMAP_FAILURE |
0 |
The other three GPUs on the same node report 0 volatile corrected errors. Across a 288-node fleet, every other GPU reporting any corrected errors at all reports 1 or 2, so this device is an outlier by roughly nine orders of magnitude.
Note there are no double-bit errors and no remap failures: ECC is correcting everything successfully. The problem appears to be the cost of doing so.
Symptom under load
A monitoring agent calling DCGM every 15s logs, repeatedly:
DCGM probe dcgm_health_check has not returned after 45.6s (deadline 45.0s)
DCGM health check timed out: Timeout. Indicating connectivity failure.
DCGM probe dcgm_cleanup has not returned after 45.6s (deadline 45.0s)
Error unwatching GPU temp limit field watch: Timeout
dcgm_health_check is the first call to stall; dcgm_cleanup stalls afterwards during the shutdown that follows. Over three days this happened 49 times, roughly 15 times per day, whenever a large distributed training job was resident on the node.
A separate full dcgmi diag run against this node also completed only "after a very long run time", and reported targeted_power Fail for this GPU alone: max power 207.4 W against a target minimum ratio of 899.2. Every other subtest, including memory, memory_bandwidth, diagnostic, nvbandwidth and pcie, passed on all four GPUs.
The same calls are instant when the GPU is idle
The node was later drained for hardware repair. With no workload resident:
| Under load | Idle | |
|---|---|---|
| Corrected ECC rate on the affected GPU | 4,647/s, then 17,657/s | 0, sustained for 2.75 days |
dcgmi health -g 0 -c |
stalls past 45s | 0s |
| Monitoring agent restarts from stalled DCGM calls | ~15/day | 0 in 2.75 days |
dcgmi discovery -l returns promptly in both states. nvidia-smi runs and reports all four GPUs in both states.
So the stall correlates with correctable-error activity on the device, not with uptime, not with DCGM's own state, and not with the driver generally.
What we ruled out
- Stale
nv-hostengine. Restarted it; a freshly started engine stalls the same way within minutes under load. The host engine process itself never crashed (0 restarts across the whole period), so it hangs while alive. - Driver or DCGM version skew. Driver 580.126.20 and kernel 6.17.0-1014-nvidia-64k are byte-identical to the node's healthy neighbours.
- A rack-wide or job-wide effect. 17 other nodes in the same rack ran ranks of the same distributed job with zero stalls and zero corrected ECC errors.
- XIDs. The node logs XID 45 only, at a volume that is mid-pack for its rack and therefore attributable to the job rather than this GPU. There are no ECC, SBE, DBE, remap or retirement messages in the kernel log at all. The only
NVRMlines are IMEX_fabricNotifyEventand a handful of NVLink status-collection failures, and all of them stopped when the node was drained.
Secondary observation: the health check reports Healthy
Possibly a separate issue, but it seems worth raising alongside:
$ dcgmi health -g 0 -s mpi
Health monitor systems set successfully.
$ dcgmi health -g 0 -c
+---------------------------+----------------------------------------------------------+
| Health Monitor Report |
+===========================+==========================================================+
| Overall Health | Healthy |
+---------------------------+----------------------------------------------------------+
Overall Health: Healthy on a GPU with 54.7 billion lifetime corrected errors and 8 correctable remapped rows. I understand the reasoning: no double-bit errors, no remap failures, and correctable errors are expected in normal operation. But there appears to be no threshold at which correctable-error volume or correctable row remapping becomes a health finding, so a device degrading this far stays Healthy until its first uncorrectable error.
If that is intended, it would help to have it documented, since consumers of the health API reasonably treat Healthy as "this GPU is fine to schedule on".
Questions
- Is a stall of this kind expected when a GPU is generating correctable ECC errors at this rate, for example because ECC state queries serialise against correction activity?
- Is there a way to bound or time-box the affected calls so a monitoring client is not blocked past its own deadline?
- Should correctable-error volume or correctable row remapping ever influence the health verdict, and if not, is there a recommended field-based signal for "this GPU is degrading" that consumers should use instead?
Possibly related but distinct: #209 describes nv-hostengine becoming unresponsive after 2-3 days of continuous running on driver 560.35.03. That one is uptime-correlated; this one is load-and-ECC-correlated and recovers fully when the GPU goes idle.
Happy to gather more detail. The node is currently idle and awaiting a hardware repair, so I can still run diagnostics against it in its degraded state, though I may not be able to reproduce the loaded condition again once the GPU is replaced.
- Dominant language
- C++
- Stars
- 798
- Forks
- 112
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/DCGM
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
dcgmi InstallCtrlHandler unhandled returnMay be free again A pull request for this issue was closed without being merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
Similar issues
-
Status: Awaiting triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
espressif/arduino-esp32#12984 ·
Maintainers usually reply within 1 day
-
torch_ops/logprob.cu does not compile with the serving container's nvcc (13.3.73); check_torch_ops.py cannot run as shippedPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 Under an hour Newbie friendliness 72/100
ashhart/TensorFold#535 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
agent:Windows bug
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day