ggml_cuda_init: failed to initialize CUDA: (null) on Windows with CUDA 12.9
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 35/100
- Issue 類型
- 缺陷
- 描述清晰度
- 基本清楚
- 活躍度
- 停滯
研究方向
這個 issue 涉及在使用 CUDA 12.9 的 Windows 上 CUDA 初始化失敗的問題。首先檢查建置記錄和 CMake 設定(GGML_CUDA 旗標)。檢查 llama.cpp 原始碼中的 ggml_cuda_init 函式,以了解 CUDA runtime 檢查。驗證 CUDA toolkit 的安裝、驅動程式相容性和環境變數。在函式庫外執行簡單的 CUDA 測試程式有助於隔離問題。
由索引模型根據 Issue 內容生成。
描述
System Information:
- OS: Windows
- GPU: NVIDIA GeForce RTX 5060 Ti
- NVIDIA Driver Version: 577.00
- CUDA Version (from
nvidia-smi): 12.9 - Python Version: 3.12
- Visual Studio: Visual Studio 2019 with "Desktop development with C++" workload
Problem Description:
I am unable to get llama-cpp-python to use my GPU. When I run a script to load a model with n_gpu_layers=-1, I get the error ggml_cuda_init: failed to initialize CUDA: (null), and all layers are
loaded on the CPU.
Troubleshooting Steps Taken:
- Installed llama-cpp-python using the following command in the "x64 Native Tools Command Prompt for VS 2019" with a Python virtual environment activated:
1 set CMAKE_ARGS="-DGGML_CUDA=on" && pip install --upgrade --force-reinstall llama-cpp-python --no-cache-dir - Verified that the command completes successfully, but the resulting installation does not use the GPU.
- Tried using the deprecated LLAMA_CUBLAS flag, which resulted in a build error (as expected).
- Performed a full cleanup of the environment:
- pip uninstall llama-cpp-python
- pip cache purge
- Manually deleted leftover ~* directories from site-packages.
- Reinstalled after the cleanup, but the problem persists.
- Installed PyTorch with CUDA 12.1 support (pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121) before reinstalling llama-cpp-python, but this did not
resolve the issue. - Confirmed that the correct Python interpreter and virtual environment are being used.
- The run_with_llama_cpp.py script being used is:
1 from llama_cpp import Llama
2
3 llm = Llama(
4 model_path="models/mistral-7b-instruct-v0.2.Q4_K_M.gguf",
5 n_gpu_layers=-1,
6 n_ctx=4096,
7 verbose=True
8 )
9
10 output = llm(
11 "AI is going to ",
12 max_tokens=32,
13 stop=["."],
14 echo=True
15 )
16
17 print(output)
Request:
Could you please provide any insights into why the CUDA initialization might be failing, or suggest any further diagnostic steps? I can provide the full verbose build log if needed.
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.5k
- 平均合併
- 3 小時 57 分鐘
- 30 天內合併 PR
- 4
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
abetlen/llama-cpp-python 的其他 Issue
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently drops可能已有人在做 @Belal0066 於 15 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
維護者通常 1 天內回覆
-
uv add llama-cpp-python wheels fails for versions above 0.3.30可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 則留言 · 2 個 reaction ·
維護者通常 1 天內回覆
-
Docs: consolidate build-from-source and GPU backend guide可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 則留言 ·
維護者通常 1 天內回覆
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusingly可能已有人在做 @Anai-Guo 於 32 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
維護者通常 1 天內回覆
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 68/100
-
Task
難度 2/5 1-3 小時 新手友好度 65/100
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 86/100
war-and-code/dircue#200 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 87/100
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 84/100
維護者通常 1 天內回覆