ggml_cuda_init: failed to initialize CUDA: (null) on Windows with CUDA 12.9
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Estancado
- Área
- ai-infra-agents, backend
Línea de trabajo
El issue trata sobre un fallo de inicialización de CUDA en Windows con CUDA 12.9. Empieza examinando los registros de compilación y la configuración de CMake (flag GGML_CUDA). Comprueba la función ggml_cuda_init en el código fuente de llama.cpp para entender las comprobaciones del runtime de CUDA. Verifica la instalación del toolkit de CUDA, la compatibilidad del controlador y las variables de entorno. Ejecutar un programa de prueba sencillo de CUDA fuera de la biblioteca puede ayudar a aislar el problema.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
System Information:
- OS: Windows
- GPU: NVIDIA GeForce RTX 5060 Ti
- NVIDIA Driver Version: 577.00
- CUDA Version (from
nvidia-smi): 12.9 - Python Version: 3.12
- Visual Studio: Visual Studio 2019 with "Desktop development with C++" workload
Problem Description:
I am unable to get llama-cpp-python to use my GPU. When I run a script to load a model with n_gpu_layers=-1, I get the error ggml_cuda_init: failed to initialize CUDA: (null), and all layers are
loaded on the CPU.
Troubleshooting Steps Taken:
- Installed llama-cpp-python using the following command in the "x64 Native Tools Command Prompt for VS 2019" with a Python virtual environment activated:
1 set CMAKE_ARGS="-DGGML_CUDA=on" && pip install --upgrade --force-reinstall llama-cpp-python --no-cache-dir - Verified that the command completes successfully, but the resulting installation does not use the GPU.
- Tried using the deprecated LLAMA_CUBLAS flag, which resulted in a build error (as expected).
- Performed a full cleanup of the environment:
- pip uninstall llama-cpp-python
- pip cache purge
- Manually deleted leftover ~* directories from site-packages.
- Reinstalled after the cleanup, but the problem persists.
- Installed PyTorch with CUDA 12.1 support (pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121) before reinstalling llama-cpp-python, but this did not
resolve the issue. - Confirmed that the correct Python interpreter and virtual environment are being used.
- The run_with_llama_cpp.py script being used is:
1 from llama_cpp import Llama
2
3 llm = Llama(
4 model_path="models/mistral-7b-instruct-v0.2.Q4_K_M.gguf",
5 n_gpu_layers=-1,
6 n_ctx=4096,
7 verbose=True
8 )
9
10 output = llm(
11 "AI is going to ",
12 max_tokens=32,
13 stop=["."],
14 echo=True
15 )
16
17 print(output)
Request:
Could you please provide any insights into why the CUDA initialization might be failing, or suggest any further diagnostic steps? I can provide the full verbose build log if needed.
- Lenguaje dominante
- Python
- Estrellas
- 10.6k
- Forks
- 1.5k
- Merge medio
- 3 h 57 min
- PR fusionados (30 d)
- 4
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsPosiblemente ocupada @Belal0066 la tomó hace 15 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
abetlen/llama-cpp-python#2371 ·
Los mantenedores suelen responder en 1 día
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Posiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2352 · 1 comentario · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
Docs: consolidate build-from-source and GPU backend guidePosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
abetlen/llama-cpp-python#2314 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2211 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyPosiblemente ocupada @Anai-Guo la tomó hace 32 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2210 ·
Los mantenedores suelen responder en 1 día
Todos los issues de abetlen/llama-cpp-python
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
-
Harmony OPeNDAP SubSetter (HOSS) Geographic LARC_CLOUD PREFIRE_SAT2_AUX-SAT R01 production
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
nasa/harmony-autotester#245 ·
-
[FEATURE] - Add UTVD supportAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Deltares/imod-python#1928 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
feature
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100