[BUG]: VirtualMemoryResource `host_numa` allocations always fail
Maintainer thường phản hồi trong vòng 1 ngày
@Andy-Jost đang làm issue này rồi.
Từ ngày 24/9/2026.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 70/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- backend, operating-systems
Hướng nghiên cứu
Bắt đầu trong cuda_core/cuda/core/_memory/_virtual_memory_resource.py, đặc biệt là VirtualMemoryResource.init và allocate, đồng thời xem xét các định nghĩa HOST_NUMA trong cuda_core/cuda/core/typing.py. Chạy cuda_core/tests/test_memory.py và mở rộng phạm vi kiểm thử vị trí host để kiểm thử việc cấp phát, bao gồm các trường hợp host_numa và host_numa_current được báo cáo. Hoàn thành khi hành vi của các vị trí host được hỗ trợ đã được kiểm thử và việc cấp phát không còn thất bại do thiếu ID vị trí.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Is this a duplicate?
- I confirmed there appear to be no duplicate issues for this bug and that I agree to the Code of Conduct
Type of Bug
Runtime Error
Component
cuda.core
Describe the bug
Every VirtualMemoryResource allocation with location_type="host_numa" fails with CUDA_ERROR_INVALID_VALUE, on a system that reports host_numa_id 0 and host_numa_virtual_memory_management_supported True. VirtualMemoryResource.allocate passes prop.location.id = -1 whenever the resource has no device, which __init__ sets for every host location type, and cuMemCreate rejects CU_MEM_LOCATION_TYPE_HOST_NUMA with that id. cuda_core/cuda/core/typing.py documents HOST_NUMA as "host memory pinned to a specific NUMA node", and VirtualMemoryResourceOptions has no field that carries a node id.
location_type="host_numa_current" also fails, for a separate reason: cuMemCreate rejects CU_MEM_LOCATION_TYPE_HOST_NUMA_CURRENT for every node id I passed it, including 0.
test_vmm_host_location_types_report_host_accessible in cuda_core/tests/test_memory.py parametrizes over host, host_numa and host_numa_current, but it only checks mr.device and mr.is_host_accessible, so it stays green without ever calling allocate.
How to Reproduce
from cuda.core import Device, VirtualMemoryResource, VirtualMemoryResourceOptions
dev = Device()
dev.set_current()
print("host_numa_id =", dev.properties.host_numa_id)
for location_type in ("host", "host_numa"):
opts = VirtualMemoryResourceOptions(location_type=location_type, handle_type=None)
buf = VirtualMemoryResource(dev, config=opts).allocate(4096)
print(location_type, "allocated", buf.size, "bytes")
buf.close()
host_numa_id = 0
host allocated 2097152 bytes
Traceback (most recent call last):
File "/tmp/vmm_host_numa.py", line 8, in <module>
buf = VirtualMemoryResource(dev, config=opts).allocate(4096)
File "/home/vyron-vasileiadis/projects/forks/cuda-python/cuda_core/cuda/core/_memory/_virtual_memory_resource.py", line 555, in allocate
raise_if_driver_error(res)
~~~~~~~~~~~~~~~~~~~~~^^^^^
File "cuda/core/_utils/cuda_utils.pyx", line 134, in cuda.core._utils.cuda_utils._check_driver_error
File "cuda/core/_utils/cuda_utils.pyx", line 145, in cuda.core._utils.cuda_utils._check_driver_error
cuda.core._utils.cuda_utils.CUDAError: CUDA_ERROR_INVALID_VALUE: This indicates that one or more of the parameters passed to the API call is not within an acceptable range of values.
Expected behavior
allocate should either return a buffer pinned to a NUMA node, the way location_type="host" and location_type="device" already do, or fail at construction with an error naming the missing node id. Driving cuMemCreate directly on this system, CU_MEM_LOCATION_TYPE_HOST_NUMA returns CUDA_SUCCESS at location.id = 0 and CUDA_ERROR_INVALID_VALUE at location.id = -1, so the location id is the only thing standing between the current behaviour and a working allocation.
Operating System
Ubuntu 26.04 LTS
nvidia-smi output
Tue Aug 25 11:22:23 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.84 Driver Version: 595.84 CUDA Version: 13.2 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 3050 ... Off | 00000000:01:00.0 Off | N/A |
| N/A 46C P8 3W / 30W | 11MiB / 4096MiB | 0% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+
- Ngôn ngữ chính
- Cython
- Star
- 3.4k
- Fork
- 334
- Merge trung bình
- 1 ngày 16 giờ
- Pull request đã merge (30 ngày)
- 127
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/cuda-python
-
triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
NVIDIA/cuda-python#2952 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[DOC]: cuda.core 1.1.1 note misstates program cache permissionsCó thể đã có người làm @leofang đã nhận 10 ngày trước. Đang mởdocumentation P1
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
NVIDIA/cuda-python#2717 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[DOC]: `PinnedMemoryResource.allocate` documents no parametersCó thể đã có người làm @Andy-Jost đã nhận 10 ngày trước. Đang mởcuda.core documentation P1
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 90/100
NVIDIA/cuda-python#2712 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[BUG]: LocatedHeaderDir is mutable, so callers can poison the cached header-directory lookupCó thể đã có người làm @rwgk đã nhận 10 ngày trước. Đang mởtriage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/cuda-python#2646 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[FEA]: Support inheritance from BufferCó thể đã có người làm @leofang đã nhận 10 ngày trước. Đang mởcuda.core triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
NVIDIA/cuda-python#2435 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA/cuda-python
Issue tương tự
-
[BUG] 请修改标题为您遇到的问题Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
OpenListTeam/OpenList-Worker#103 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
needs-ac
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Ikalus1988/MisakaNet#2845 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
modelcontextprotocol/go-sdk#1340 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
skills.mdx: ReadResourceDirectoryRequest does not type-check against the 2026-07-28 base schemaĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
modelcontextprotocol/ext-skills#156 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
MystenLabs/MemWal#1104 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày