Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

ggml_cuda_init: failed to initialize CUDA: (null) on Windows with CUDA 12.9

Aberta
#2,062 2 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
4/5
Tempo estimado
3-5 dias
Facilidade para iniciantes
35/100
Tipo de issue
Bug
Clareza
Razoavelmente clara
Status de atividade
Estagnada
Stack de tecnologia
c, cmake, python

Direção de pesquisa

A issue trata de uma falha na inicialização do CUDA no Windows com CUDA 12.9. Comece examinando os logs de build e a configuração do CMake (flag GGML_CUDA). Verifique a função ggml_cuda_init no código-fonte do llama.cpp para entender as verificações do runtime do CUDA. Verifique a instalação do toolkit do CUDA, a compatibilidade do driver e as variáveis de ambiente. Executar um programa de teste simples de CUDA fora da biblioteca pode ajudar a isolar o problema.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

System Information:

  • OS: Windows
  • GPU: NVIDIA GeForce RTX 5060 Ti
  • NVIDIA Driver Version: 577.00
  • CUDA Version (from nvidia-smi): 12.9
  • Python Version: 3.12
  • Visual Studio: Visual Studio 2019 with "Desktop development with C++" workload

Problem Description:
I am unable to get llama-cpp-python to use my GPU. When I run a script to load a model with n_gpu_layers=-1, I get the error ggml_cuda_init: failed to initialize CUDA: (null), and all layers are
loaded on the CPU.

Troubleshooting Steps Taken:

  1. Installed llama-cpp-python using the following command in the "x64 Native Tools Command Prompt for VS 2019" with a Python virtual environment activated:
    1 set CMAKE_ARGS="-DGGML_CUDA=on" && pip install --upgrade --force-reinstall llama-cpp-python --no-cache-dir
  2. Verified that the command completes successfully, but the resulting installation does not use the GPU.
  3. Tried using the deprecated LLAMA_CUBLAS flag, which resulted in a build error (as expected).
  4. Performed a full cleanup of the environment:
    • pip uninstall llama-cpp-python
    • pip cache purge
    • Manually deleted leftover ~* directories from site-packages.
  5. Reinstalled after the cleanup, but the problem persists.
  6. Installed PyTorch with CUDA 12.1 support (pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121) before reinstalling llama-cpp-python, but this did not
    resolve the issue.
  7. Confirmed that the correct Python interpreter and virtual environment are being used.
  8. The run_with_llama_cpp.py script being used is:
1     from llama_cpp import Llama
2 
3     llm = Llama(
4       model_path="models/mistral-7b-instruct-v0.2.Q4_K_M.gguf",
5       n_gpu_layers=-1,
6       n_ctx=4096,
7       verbose=True
8     )
9 

10 output = llm(
11 "AI is going to ",
12 max_tokens=32,
13 stop=["."],
14 echo=True
15 )
16
17 print(output)

Request:
Could you please provide any insights into why the CUDA initialization might be failing, or suggest any further diagnostic steps? I can provide the full verbose build log if needed.

Linguagem predominante
Python
Estrelas
10.6k
Forks
1.5k
Merge médio
3h 57min
PRs com merge (30d)
4

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de abetlen/llama-cpp-python

Todas as issues de abetlen/llama-cpp-python

Issues semelhantes

Mais issues de Python

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.