Improve installation DX: prebuilt wheels for 3.13/3.14/3.14t + declarative backend selection
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 3/5
- Temps estimé
- 1-2 jours
- Accessibilité débutants
- 45/100
- Type d'issue
- Fonctionnalité
- Clarté
- Clairement spécifiée
- Activité
- À l'abandon
- Stack technique
- cmake, python
- Domaine
- build-system, ci-cd, developer-experience, tooling
Piste de recherche
L’issue pointe vers trois fichiers spécifiques de workflow CI : .github/workflows/build-wheels-metal.yaml, build-wheels-cuda.yaml et build-and-release.yaml. Commencez par examiner les versions actuelles de cibuildwheel et les matrices CIBW_BUILD dans ces fichiers. La modification consiste à mettre à jour la version et à ajouter la prise en charge des builds Python 3.13, 3.14 et free-threaded. Testez les changements localement ou dans un fork afin de vérifier que les builds de wheels réussissent. « Done » signifie que les workflows produisent des wheels pour les nouvelles versions de Python et que la documentation est mise à jour pour mentionner l’option config-settings -C cmake.args.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Problem
Installing llama-cpp-python with a GPU backend requires setting CMAKE_ARGS as an environment variable at build time:
CMAKE_ARGS="-DGGML_METAL=on" pip install llama-cpp-python
This creates pain across the ecosystem:
-
Not declarable in
pyproject.toml— Every downstream project needs custom Makefiles or install scripts with GPU auto-detection logic (macOS → Metal, nvidia-smi → CUDA, rocminfo → ROCm, fallback → OpenBLAS). This is duplicated across hundreds of projects. -
Cache invalidation is broken —
pipanduvcache wheels by package version, not byCMAKE_ARGS. A cached OpenBLAS wheel silently gets reused when Metal or CUDA is requested. Workaround:--no-cache, which defeats caching entirely. -
GPU prebuilt wheels stop at Python 3.12 — The Metal wheel CI (
build-wheels-metal.yaml) is hardcoded toCIBW_BUILD: "cp39-* cp310-* cp311-* cp312-*". The CUDA wheel CI (build-wheels-cuda.yaml) has its matrix pinned to Python 3.9-3.12. CPU-only wheels include 3.13 (via default cibuildwheel config inbuild-and-release.yaml), but the arm64 job there also pins to cp38-cp312. No workflow produces 3.14 or free-threaded (3.13t/3.14t) wheels. Python 3.13 has been stable since Oct 2024, 3.14 since Oct 2025. Free-threaded builds are increasingly important — vLLM, llguidance, and the broader no-GIL ecosystem depend on them.
Current state of published wheel indexes:
| Index | cp313 | cp314 | Free-threaded |
|---|---|---|---|
CPU (/whl/cpu/) |
✅ | ❌ | ❌ |
Metal (/whl/metal/) |
❌ | ❌ | ❌ |
CUDA (/whl/cu1xx/) |
❌ | ❌ | ❌ |
Proposed changes
1. Expand prebuilt wheel matrix (highest impact, smallest change)
Update CIBW_BUILD in Metal/CUDA workflows and add free-threaded support. This is the single highest-impact change — it eliminates source builds for most users.
build-wheels-metal.yaml:
Upgrade cibuildwheel from v2.22.0 to v3.x (3.0 added cp314/cp314t support). In cibuildwheel 3.0, cp314t is built by default (free-threading is no longer experimental in 3.14), and cp313t requires CIBW_ENABLE: cpython-freethreading.
- uses: pypa/[email protected]
+ uses: pypa/[email protected]
env:
- CIBW_BUILD: "cp39-* cp310-* cp311-* cp312-*"
+ CIBW_BUILD: "cp39-* cp310-* cp311-* cp312-* cp313-* cp314-*"
+ CIBW_ENABLE: cpython-freethreading
build-and-release.yaml — same cibuildwheel upgrade, and update the build_wheels_arm64 job:
- CIBW_BUILD: "cp38-* cp39-* cp310-* cp311-* cp312-*"
+ CIBW_BUILD: "cp38-* cp39-* cp310-* cp311-* cp312-* cp313-* cp314-*"
+ CIBW_ENABLE: cpython-freethreading
build-wheels-cuda.yaml — uses a different build system (python -m build --wheel with a PowerShell matrix). The pyver matrix would need "3.13", "3.14" added.
With prebuilt wheels, any downstream project can use uv's declarative index support:
# pyproject.toml — zero Makefile, zero CMAKE_ARGS
[project]
dependencies = ["llama-cpp-python~=0.3"]
[tool.uv.sources]
llama-cpp-python = [
{ index = "llama-metal", marker = "sys_platform == 'darwin'" },
{ index = "llama-cpu", marker = "sys_platform == 'linux'" },
]
[[tool.uv.index]]
name = "llama-metal"
url = "https://abetlen.github.io/llama-cpp-python/whl/metal"
explicit = true
[[tool.uv.index]]
name = "llama-cpu"
url = "https://abetlen.github.io/llama-cpp-python/whl/cpu"
explicit = true
2. Document --config-settings as the source-build path
Since the build backend is scikit-build-core, cmake args can be passed via the standard PEP 517 config-settings interface:
pip install llama-cpp-python -C cmake.args="-DGGML_METAL=on"
# or with uv:
uv pip install llama-cpp-python -C cmake.args="-DGGML_METAL=on"
This is cleaner than the CMAKE_ARGS env var — it's the standard PEP 517 mechanism, more explicit, and discoverable. It's already supported via scikit-build-core but not documented in the README or install docs.
3. (Future) Adopt PEP 817 Wheel Variants
PEP 817 (draft, Dec 2025) introduces a standard mechanism for GPU/accelerator wheel variants. PyTorch 2.9 already ships experimental variant-enabled wheels. Once PEP 817 is accepted and tool support lands, llama-cpp-python could publish variant wheels that are auto-selected by the installer:
# Future: just works, installer picks Metal/CUDA/CPU automatically
pip install llama-cpp-python
This is mentioned for context only — the actionable items are (1) and (2) above.
Ecosystem context
- Quansight offered funded engineering help for free-threaded support in #2103 (via vLLM ecosystem work) — awaiting maintainer signal
- ~470K monthly PyPI downloads (pypistats) — every project using this beyond toy scripts hits this install wall
- How others solved it: PyTorch uses per-backend index URLs + PEP 817 variants; ONNX Runtime publishes separate PyPI packages per backend (
onnxruntime-gpu,onnxruntime-silicon)
Related
Wheel matrix gaps (same root cause):
- #2103 — Pre-built wheels for Python 3.14 and 3.14 free-threaded
- #2130 — Pre-built CPU-only wheel for Windows (cp313)
- #2068 — Where can I download wheel for CUDA 12.8?
- #2091 — CUDA 12.8 wheel request
Wheel variants / long-term packaging:
- #2092 — Add support for experimental wheel variants (wheelnext)
- #1506 — Multi-arch support for pre-built CPU wheel (by @abetlen)
- Discussion #1875 — Automating pre-building of wheels for all platforms
Downstream impact of missing wheels:
- #2118 — Installation deadlock on Hugging Face Spaces (musl/glibc mismatch)
- #2113 — No working wheels for Debian/Ubuntu
Happy to submit a PR for (1) and (2).
- Langage dominant
- Python
- Étoiles
- 10.6k
- Forks
- 1.5k
- Merge moyen
- 3 h 57 min
- PR mergées (30 j)
- 4
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Aucun modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsPeut-être pris @Belal0066 l’a pris il y a 17 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 88/100
abetlen/llama-cpp-python#2371 ·
Les mainteneurs répondent en général sous 1 jour
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Peut-être pris Une pull request liée à cette issue est ouverte ou déjà fusionnée. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2352 · 1 commentaire · 2 réactions ·
Les mainteneurs répondent en général sous 1 jour
-
Docs: consolidate build-from-source and GPU backend guidePeut-être pris Une pull request liée à cette issue est ouverte ou déjà fusionnée. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
abetlen/llama-cpp-python#2314 ·
Les mainteneurs répondent en général sous 1 jour
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayPeut-être à nouveau libre @lxcxjxhx l’a pris il y a 94 jours, et aucune pull request n’est ouverte. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2211 · 2 commentaires ·
Les mainteneurs répondent en général sous 1 jour
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyPeut-être pris @Anai-Guo l’a pris il y a 34 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2210 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de abetlen/llama-cpp-python
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
UKGovernmentBEIS/inspect_ai#5802 ·
Les mainteneurs répondent en général sous 2 jours
-
Difficulté 2/5 1-3 heures Accessibilité débutants 74/100
no-human-ai/no_human#660 ·
Les mainteneurs répondent en général sous 1 jour
-
Add a Security Insights v2 fileOuvertedocumentation good first issue
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
Les mainteneurs répondent en général sous 1 jour
-
documentation need help question
Difficulté 1/5 1-3 heures Accessibilité débutants 66/100
phonology024/babelscribe#26 ·
-
bug
Difficulté 2/5 1-3 heures Accessibilité débutants 62/100