ggml_cuda_init: failed to initialize CUDA: (null) on Windows with CUDA 12.9
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 35/100
- Tipo de issue
- Bug
- Clareza
- Razoavelmente clara
- Status de atividade
- Estagnada
- Domínio
- ai-infra-agents, backend
Direção de pesquisa
A issue trata de uma falha na inicialização do CUDA no Windows com CUDA 12.9. Comece examinando os logs de build e a configuração do CMake (flag GGML_CUDA). Verifique a função ggml_cuda_init no código-fonte do llama.cpp para entender as verificações do runtime do CUDA. Verifique a instalação do toolkit do CUDA, a compatibilidade do driver e as variáveis de ambiente. Executar um programa de teste simples de CUDA fora da biblioteca pode ajudar a isolar o problema.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
System Information:
- OS: Windows
- GPU: NVIDIA GeForce RTX 5060 Ti
- NVIDIA Driver Version: 577.00
- CUDA Version (from
nvidia-smi): 12.9 - Python Version: 3.12
- Visual Studio: Visual Studio 2019 with "Desktop development with C++" workload
Problem Description:
I am unable to get llama-cpp-python to use my GPU. When I run a script to load a model with n_gpu_layers=-1, I get the error ggml_cuda_init: failed to initialize CUDA: (null), and all layers are
loaded on the CPU.
Troubleshooting Steps Taken:
- Installed llama-cpp-python using the following command in the "x64 Native Tools Command Prompt for VS 2019" with a Python virtual environment activated:
1 set CMAKE_ARGS="-DGGML_CUDA=on" && pip install --upgrade --force-reinstall llama-cpp-python --no-cache-dir - Verified that the command completes successfully, but the resulting installation does not use the GPU.
- Tried using the deprecated LLAMA_CUBLAS flag, which resulted in a build error (as expected).
- Performed a full cleanup of the environment:
- pip uninstall llama-cpp-python
- pip cache purge
- Manually deleted leftover ~* directories from site-packages.
- Reinstalled after the cleanup, but the problem persists.
- Installed PyTorch with CUDA 12.1 support (pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121) before reinstalling llama-cpp-python, but this did not
resolve the issue. - Confirmed that the correct Python interpreter and virtual environment are being used.
- The run_with_llama_cpp.py script being used is:
1 from llama_cpp import Llama
2
3 llm = Llama(
4 model_path="models/mistral-7b-instruct-v0.2.Q4_K_M.gguf",
5 n_gpu_layers=-1,
6 n_ctx=4096,
7 verbose=True
8 )
9
10 output = llm(
11 "AI is going to ",
12 max_tokens=32,
13 stop=["."],
14 echo=True
15 )
16
17 print(output)
Request:
Could you please provide any insights into why the CUDA initialization might be failing, or suggest any further diagnostic steps? I can provide the full verbose build log if needed.
- Linguagem predominante
- Python
- Estrelas
- 10.6k
- Forks
- 1.5k
- Merge médio
- 3h 57min
- PRs com merge (30d)
- 4
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Sem modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsTalvez já em andamento @Belal0066 assumiu há 15 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
abetlen/llama-cpp-python#2371 ·
Mantenedores costumam responder em até 1 dia
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Talvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2352 · 1 comentário · 2 reações ·
Mantenedores costumam responder em até 1 dia
-
Docs: consolidate build-from-source and GPU backend guideTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
abetlen/llama-cpp-python#2314 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2211 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyTalvez já em andamento @Anai-Guo assumiu há 32 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2210 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de abetlen/llama-cpp-python
Issues semelhantes
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 92/100
-
Harmony OPeNDAP SubSetter (HOSS) Geographic LARC_CLOUD PREFIRE_SAT2_AUX-SAT R01 production
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
nasa/harmony-autotester#245 ·
-
[FEATURE] - Add UTVD supportAbertaenhancement
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Deltares/imod-python#1928 ·
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
Mantenedores costumam responder em até 1 dia
-
feature
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 66/100