Expose `ggml_backend_load()` and `ggml_backend_load_all()` to make use of builds with `GGML_BACKEND_DL=ON` and `GGML_CPU_ALL_VARIANTS=ON`
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 3/5
- Tempo estimado
- 1-2 dias
- Facilidade para iniciantes
- 45/100
- Tipo de issue
- Funcionalidade
- Clareza
- Razoavelmente clara
- Status de atividade
- Estagnada
- Domínio
- backend, build-system, performance
Direção de pesquisa
Examine os bindings C existentes em llama-cpp-python para funções relacionadas ao backend. A issue menciona ggml_backend_load() e ggml_backend_load_all() do llama.cpp; encontre onde funções semelhantes são expostas e adicione-as. Verifique a configuração de build com GGML_BACKEND_DL=ON para entender o mecanismo de carregamento dinâmico. Teste compilando e carregando um modelo para garantir que as bibliotecas do backend sejam carregadas corretamente e que o erro seja resolvido.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
I just tried compiling llama-cpp-python with GGML_BACKEND_DL=ON and GGML_CPU_ALL_VARIANTS=ON to make use of this nice feature with dynamic dispatch to a dynamically loaded backend, which e.g. made it possible to build llama.cpp once but dynamically choose the best backend for the current CPU, i.e. for x86_64 depending on whether certain instructions like AVX2 or AVX512 are available choose the best backend for the current microarchitecture level.
Compiling worked for me so far on Ubuntu 24.04 LTS and when inspecting the wheel I see the backend dynamic libraries like bin/libggml-cpu-x64.so, libggml-cpu-sse42.so, libggml-cpu-haswell.so and so on. So that is good already.
But when loading a model with llama-cpp-python I get this error:
llama_model_load_from_file_impl: no backends are loaded. hint: use ggml_backend_load() or ggml_backend_load_all() to load a backend before calling this function but these functions are not exposed yet via the bindings.
I think this would be a really great thing to add. That would make the CPU wheels for llama-cpp-python way better, because it wouldn't be stuck with base x86_64 instructions and could thus be way more performant for cases where the wheel cannot be compiled at installation time.
- Linguagem predominante
- Python
- Estrelas
- 10.6k
- Forks
- 1.5k
- Merge médio
- 3h 57min
- PRs com merge (30d)
- 4
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Sem modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsTalvez já em andamento @Belal0066 assumiu há 18 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
abetlen/llama-cpp-python#2371 ·
Mantenedores costumam responder em até 1 dia
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Talvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2352 · 1 comentário · 2 reações ·
Mantenedores costumam responder em até 1 dia
-
Docs: consolidate build-from-source and GPU backend guideTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
abetlen/llama-cpp-python#2314 ·
Mantenedores costumam responder em até 1 dia
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayTalvez livre de novo @lxcxjxhx assumiu há 95 dias e não há nenhum pull request aberto. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2211 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyTalvez já em andamento @Anai-Guo assumiu há 36 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2210 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de abetlen/llama-cpp-python
Issues semelhantes
-
enhancement good first issue
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
-
python-version
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
Mantenedores costumam responder em até 1 dia
-
bug javascript P2-medium python release:v3.1
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
adrirubio/claude-deck#546 ·
Mantenedores costumam responder em até 1 dia
-
area: desktop area: website priority: P2 type: feature
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
appandflow/stim#3411 · 1 comentário ·
Mantenedores costumam responder em até 1 dia