Add support for experimental wheel variants (i.e., wheelnext)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- python
- Ambito
- build-system, tooling
Direzione di ricerca
La issue riguarda la modifica del processo di build e pubblicazione delle wheel. Inizia esaminando gli script di build del progetto, probabilmente in setup.py o pyproject.toml, e i workflow CI/CD. Studia la specifica WheelNext e il modo in cui progetti come PyTorch implementano i metadati delle varianti. L'obiettivo è produrre wheel specifiche per il backend (CUDA, ROCm, Metal) con i metadati corretti, assicurandosi che le wheel CPU rimangano disponibili come fallback. I test consisteranno nella compilazione locale delle wheel e nella verifica dei metadati.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Is your feature request related to a problem? Please describe.
Today, installing llama-cpp-python on machines with different GPU backends (CUDA, ROCm, Metal, etc.) requires separate package names, custom extra indexes, or installer-level logic to select the correct wheel. This creates friction for downstream tooling (CLIs, orchestrators, and packaging systems) that want to provide a “just works” experience, especially when users don’t know which backend they need. Even a simple developer-driven install might require picking precisely the correct wheel.
Describe the solution you'd like
Add support for WheelNext-compatible experimental wheel variants when building and publishing wheels.
This would allow llama-cpp-python to produce a single package version that provides multiple backend-aware binary wheels, each annotated with variant metadata (e.g., GPU type, CUDA version, ROCm version).
Installers that understand the WheelNext spec (now used experimentally by PyTorch, uv, and others) can automatically select the correct backend wheel based on the system’s hardware/software configuration without a need for custom index URLs, separate packages, or manual backend flags.
Key pieces:
- Generate wheels with variant metadata following the experimental WheelNext (wheel variants) conventions.
- Publish per-backend wheels using the standardized naming + metadata fields.
- Ensure that CPU-only wheels remain available as fallback.
This would significantly simplify installation for all users and remove backend-selection logic from downstream tools. Wheel variants are fully backward-compatible so existing workflows won't be disrupted.
Describe alternatives you've considered
- Separate package names per backend (e.g., llama-cpp-python-cuda): fragments packaging and forces manual selection.
- Extras for backend variants (pip install llama-cpp-python[cuda]): still requires external detection and doesn’t integrate with hardware-aware installer selection.
- Custom index URLs for backend wheels: brittle and requires orchestration logic outside Python packaging.
- CLI-backed installation routing (what many downstream projects do currently): it’s reinventing the wheel and provides an inconsistent experience for end users.
All of these solutions put the burden on downstream tooling rather than on standardized wheel metadata.
Additional context
- https://wheelnext.dev
- https://pytorch.org/blog/pytorch-wheel-variants/
- https://astral.sh/blog/wheel-variants
- https://labs.quansight.org/blog/python-wheels-from-tags-to-variants
- https://developer.nvidia.com/blog/streamline-cuda-accelerated-python-install-and-packaging-workflows-with-wheel-variants
- https://lwn.net/Articles/1028299/
- Lingua principale
- Python
- Stelle
- 10.6k
- Fork
- 1.5k
- Merge medio
- 3h 57m
- PR unite (30g)
- 4
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsForse già presa @Belal0066 l’ha presa 18 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
abetlen/llama-cpp-python#2371 ·
I maintainer di solito rispondono entro 1 giorno
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Forse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
abetlen/llama-cpp-python#2352 · 1 commento · 2 reazioni ·
I maintainer di solito rispondono entro 1 giorno
-
Docs: consolidate build-from-source and GPU backend guideForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
abetlen/llama-cpp-python#2314 ·
I maintainer di solito rispondono entro 1 giorno
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayForse di nuovo libera @lxcxjxhx l’ha presa 94 giorni fa e non c’è nessuna pull request aperta. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
abetlen/llama-cpp-python#2211 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyForse già presa @Anai-Guo l’ha presa 35 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
abetlen/llama-cpp-python#2210 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di abetlen/llama-cpp-python
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
aicell-lab/bioengine#232 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
modelscope/evalscope#1836 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
bug
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
jbaruch/speaker-toolkit#480 ·
I maintainer di solito rispondono entro 1 giorno