Misleading field label "GPU Device IDs Detected" in dcgmi diag output
Nobody has claimed this yet.
- #286 by @cluster2600 — closed without merging
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
Research direction
Start by running dcgmi diag --run 1 on a system with multiple identical GPUs and trace where the diagnostic metadata label and values are produced. Resolve whether the output should rename the field to describe PCI/model IDs or report unique identifiers, then verify the result against the reproduced multi-GPU output.
Written by the indexing model from the issue text.
Description
DCGM Version: 4.4.1 / 4.5.2
Description:
The dcgmi diag output displays a field labeled "GPU Device IDs Detected" which shows identical values for all GPUs:
| GPU Device IDs Detected | 3182, 3182, 3182, 3182, 3182, 3182, 3182 |
This is misleading because:
- "GPU Device IDs" implies unique per-GPU identifiers
- Users expect N different values for N GPUs
- The actual values shown are PCI Device IDs (hardware SKU), which are identical for all GPUs of the same model
Expected behavior:
Either:
- Rename the field to accurately reflect what it shows: "PCI Device IDs Detected" or "GPU Model IDs"
- Or show actual unique GPU identifiers (UUIDs, indices, or serial numbers)
Reproduction:
bash
dcgmi diag --run 1
On any system with multiple identical GPUs.
Repro stdout examples:
Successfully ran diagnostic for group.
+---------------------------+------------------------------------------------+
| Diagnostic | Result |
+===========================+================================================+
|----- Metadata ----------+------------------------------------------------|
| DCGM Version | 4.4.1 |
| Driver Version Detected | 580.95.05 |
| GPU Device IDs Detected | 3182, 3182, 3182, 3182, 3182, 3182, 3182, 3182 |
|----- Deployment --------+------------------------------------------------|
| software | Pass |
| | GPU0: Pass |
| | GPU1: Pass |
| | GPU2: Pass |
| | GPU3: Pass |
| | GPU4: Pass |
| | GPU5: Pass |
| | GPU6: Pass |
| | GPU7: Pass |
+---------------------------+------------------------------------------------+
Successfully ran diagnostic for group.
+---------------------------+------------------------------------------------+
| Diagnostic | Result |
+===========================+================================================+
|----- Metadata ----------+------------------------------------------------|
| DCGM Version | 4.5.2 |
| Driver Version Detected | 580.126.09 |
| GPU Device IDs Detected | 3182, 3182, 3182, 3182, 3182, 3182, 3182 |
|----- Deployment --------+------------------------------------------------|
| software | Pass |
| | GPU0: Pass |
| | GPU1: Pass |
| | GPU2: Pass |
| | GPU3: Pass |
| | GPU4: Pass |
| | GPU5: Pass |
| | GPU6: Pass |
+---------------------------+------------------------------------------------+
- Dominant language
- C++
- Stars
- 798
- Forks
- 112
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/DCGM
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
dcgmi InstallCtrlHandler unhandled returnMay be free again A pull request for this issue was closed without being merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
Similar issues
-
Status: Awaiting triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
espressif/arduino-esp32#12984 ·
Maintainers usually reply within 1 day
-
torch_ops/logprob.cu does not compile with the serving container's nvcc (13.3.73); check_torch_ops.py cannot run as shippedPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 Under an hour Newbie friendliness 72/100
ashhart/TensorFold#535 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day