Improve installation DX: prebuilt wheels for 3.13/3.14/3.14t + declarative backend selection
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- cmake, python
- Lĩnh vực
- build-system, ci-cd, developer-experience, tooling
Hướng nghiên cứu
Issue này chỉ ra ba tệp CI workflow cụ thể: .github/workflows/build-wheels-metal.yaml, build-wheels-cuda.yaml và build-and-release.yaml. Trước tiên, hãy kiểm tra các phiên bản cibuildwheel hiện tại và các ma trận CIBW_BUILD trong những tệp đó. Thay đổi này bao gồm việc cập nhật phiên bản và thêm hỗ trợ build cho Python 3.13, 3.14 và free-threaded. Hãy kiểm thử các thay đổi cục bộ hoặc trong một fork để đảm bảo các bản build wheel thành công. “Done” nghĩa là các workflow tạo ra wheel cho những phiên bản Python mới và tài liệu được cập nhật để đề cập đến tùy chọn config-settings -C cmake.args.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Problem
Installing llama-cpp-python with a GPU backend requires setting CMAKE_ARGS as an environment variable at build time:
CMAKE_ARGS="-DGGML_METAL=on" pip install llama-cpp-python
This creates pain across the ecosystem:
-
Not declarable in
pyproject.toml— Every downstream project needs custom Makefiles or install scripts with GPU auto-detection logic (macOS → Metal, nvidia-smi → CUDA, rocminfo → ROCm, fallback → OpenBLAS). This is duplicated across hundreds of projects. -
Cache invalidation is broken —
pipanduvcache wheels by package version, not byCMAKE_ARGS. A cached OpenBLAS wheel silently gets reused when Metal or CUDA is requested. Workaround:--no-cache, which defeats caching entirely. -
GPU prebuilt wheels stop at Python 3.12 — The Metal wheel CI (
build-wheels-metal.yaml) is hardcoded toCIBW_BUILD: "cp39-* cp310-* cp311-* cp312-*". The CUDA wheel CI (build-wheels-cuda.yaml) has its matrix pinned to Python 3.9-3.12. CPU-only wheels include 3.13 (via default cibuildwheel config inbuild-and-release.yaml), but the arm64 job there also pins to cp38-cp312. No workflow produces 3.14 or free-threaded (3.13t/3.14t) wheels. Python 3.13 has been stable since Oct 2024, 3.14 since Oct 2025. Free-threaded builds are increasingly important — vLLM, llguidance, and the broader no-GIL ecosystem depend on them.
Current state of published wheel indexes:
| Index | cp313 | cp314 | Free-threaded |
|---|---|---|---|
CPU (/whl/cpu/) |
✅ | ❌ | ❌ |
Metal (/whl/metal/) |
❌ | ❌ | ❌ |
CUDA (/whl/cu1xx/) |
❌ | ❌ | ❌ |
Proposed changes
1. Expand prebuilt wheel matrix (highest impact, smallest change)
Update CIBW_BUILD in Metal/CUDA workflows and add free-threaded support. This is the single highest-impact change — it eliminates source builds for most users.
build-wheels-metal.yaml:
Upgrade cibuildwheel from v2.22.0 to v3.x (3.0 added cp314/cp314t support). In cibuildwheel 3.0, cp314t is built by default (free-threading is no longer experimental in 3.14), and cp313t requires CIBW_ENABLE: cpython-freethreading.
- uses: pypa/[email protected]
+ uses: pypa/[email protected]
env:
- CIBW_BUILD: "cp39-* cp310-* cp311-* cp312-*"
+ CIBW_BUILD: "cp39-* cp310-* cp311-* cp312-* cp313-* cp314-*"
+ CIBW_ENABLE: cpython-freethreading
build-and-release.yaml — same cibuildwheel upgrade, and update the build_wheels_arm64 job:
- CIBW_BUILD: "cp38-* cp39-* cp310-* cp311-* cp312-*"
+ CIBW_BUILD: "cp38-* cp39-* cp310-* cp311-* cp312-* cp313-* cp314-*"
+ CIBW_ENABLE: cpython-freethreading
build-wheels-cuda.yaml — uses a different build system (python -m build --wheel with a PowerShell matrix). The pyver matrix would need "3.13", "3.14" added.
With prebuilt wheels, any downstream project can use uv's declarative index support:
# pyproject.toml — zero Makefile, zero CMAKE_ARGS
[project]
dependencies = ["llama-cpp-python~=0.3"]
[tool.uv.sources]
llama-cpp-python = [
{ index = "llama-metal", marker = "sys_platform == 'darwin'" },
{ index = "llama-cpu", marker = "sys_platform == 'linux'" },
]
[[tool.uv.index]]
name = "llama-metal"
url = "https://abetlen.github.io/llama-cpp-python/whl/metal"
explicit = true
[[tool.uv.index]]
name = "llama-cpu"
url = "https://abetlen.github.io/llama-cpp-python/whl/cpu"
explicit = true
2. Document --config-settings as the source-build path
Since the build backend is scikit-build-core, cmake args can be passed via the standard PEP 517 config-settings interface:
pip install llama-cpp-python -C cmake.args="-DGGML_METAL=on"
# or with uv:
uv pip install llama-cpp-python -C cmake.args="-DGGML_METAL=on"
This is cleaner than the CMAKE_ARGS env var — it's the standard PEP 517 mechanism, more explicit, and discoverable. It's already supported via scikit-build-core but not documented in the README or install docs.
3. (Future) Adopt PEP 817 Wheel Variants
PEP 817 (draft, Dec 2025) introduces a standard mechanism for GPU/accelerator wheel variants. PyTorch 2.9 already ships experimental variant-enabled wheels. Once PEP 817 is accepted and tool support lands, llama-cpp-python could publish variant wheels that are auto-selected by the installer:
# Future: just works, installer picks Metal/CUDA/CPU automatically
pip install llama-cpp-python
This is mentioned for context only — the actionable items are (1) and (2) above.
Ecosystem context
- Quansight offered funded engineering help for free-threaded support in #2103 (via vLLM ecosystem work) — awaiting maintainer signal
- ~470K monthly PyPI downloads (pypistats) — every project using this beyond toy scripts hits this install wall
- How others solved it: PyTorch uses per-backend index URLs + PEP 817 variants; ONNX Runtime publishes separate PyPI packages per backend (
onnxruntime-gpu,onnxruntime-silicon)
Related
Wheel matrix gaps (same root cause):
- #2103 — Pre-built wheels for Python 3.14 and 3.14 free-threaded
- #2130 — Pre-built CPU-only wheel for Windows (cp313)
- #2068 — Where can I download wheel for CUDA 12.8?
- #2091 — CUDA 12.8 wheel request
Wheel variants / long-term packaging:
- #2092 — Add support for experimental wheel variants (wheelnext)
- #1506 — Multi-arch support for pre-built CPU wheel (by @abetlen)
- Discussion #1875 — Automating pre-building of wheels for all platforms
Downstream impact of missing wheels:
- #2118 — Installation deadlock on Hugging Face Spaces (musl/glibc mismatch)
- #2113 — No working wheels for Debian/Ubuntu
Happy to submit a PR for (1) and (2).
- Ngôn ngữ chính
- Python
- Star
- 10.6k
- Fork
- 1.5k
- Merge trung bình
- 3 giờ 57 phút
- Pull request đã merge (30 ngày)
- 4
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsCó thể đã có người làm @Belal0066 đã nhận 15 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
abetlen/llama-cpp-python#2371 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2352 · 1 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Docs: consolidate build-from-source and GPU backend guideCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
abetlen/llama-cpp-python#2314 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayCó thể làm lại được @lxcxjxhx đã nhận 92 ngày trước và không có pull request nào đang mở. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2211 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyCó thể đã có người làm @Anai-Guo đã nhận 33 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2210 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của abetlen/llama-cpp-python
Issue tương tự
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 85/100
pytest-dev/pluggy#757 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 85/100
NousResearch/hermes-agent#134960 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
HTML backend: `<br>` leaks the internal sentinel U+E000 into list items, headings and captionsCó thể đã có người làm @morten-lagabote đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
docling-project/docling#4671 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
good first issue hacktoberfest infra
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày