Add support for experimental wheel variants (i.e., wheelnext)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- build-system, tooling
Hướng nghiên cứu
Issue này nói về việc sửa đổi quy trình build và phát hành wheel. Trước tiên, hãy xem xét các script build của dự án, có thể nằm trong setup.py hoặc pyproject.toml, cùng các workflow CI/CD. Nghiên cứu đặc tả WheelNext và cách các dự án như PyTorch triển khai metadata biến thể. Mục tiêu là tạo ra các wheel dành riêng cho backend (CUDA, ROCm, Metal) với metadata chính xác, đồng thời đảm bảo các wheel CPU vẫn là fallback. Việc kiểm thử sẽ bao gồm build wheel cục bộ và xác minh metadata.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Is your feature request related to a problem? Please describe.
Today, installing llama-cpp-python on machines with different GPU backends (CUDA, ROCm, Metal, etc.) requires separate package names, custom extra indexes, or installer-level logic to select the correct wheel. This creates friction for downstream tooling (CLIs, orchestrators, and packaging systems) that want to provide a “just works” experience, especially when users don’t know which backend they need. Even a simple developer-driven install might require picking precisely the correct wheel.
Describe the solution you'd like
Add support for WheelNext-compatible experimental wheel variants when building and publishing wheels.
This would allow llama-cpp-python to produce a single package version that provides multiple backend-aware binary wheels, each annotated with variant metadata (e.g., GPU type, CUDA version, ROCm version).
Installers that understand the WheelNext spec (now used experimentally by PyTorch, uv, and others) can automatically select the correct backend wheel based on the system’s hardware/software configuration without a need for custom index URLs, separate packages, or manual backend flags.
Key pieces:
- Generate wheels with variant metadata following the experimental WheelNext (wheel variants) conventions.
- Publish per-backend wheels using the standardized naming + metadata fields.
- Ensure that CPU-only wheels remain available as fallback.
This would significantly simplify installation for all users and remove backend-selection logic from downstream tools. Wheel variants are fully backward-compatible so existing workflows won't be disrupted.
Describe alternatives you've considered
- Separate package names per backend (e.g., llama-cpp-python-cuda): fragments packaging and forces manual selection.
- Extras for backend variants (pip install llama-cpp-python[cuda]): still requires external detection and doesn’t integrate with hardware-aware installer selection.
- Custom index URLs for backend wheels: brittle and requires orchestration logic outside Python packaging.
- CLI-backed installation routing (what many downstream projects do currently): it’s reinventing the wheel and provides an inconsistent experience for end users.
All of these solutions put the burden on downstream tooling rather than on standardized wheel metadata.
Additional context
- https://wheelnext.dev
- https://pytorch.org/blog/pytorch-wheel-variants/
- https://astral.sh/blog/wheel-variants
- https://labs.quansight.org/blog/python-wheels-from-tags-to-variants
- https://developer.nvidia.com/blog/streamline-cuda-accelerated-python-install-and-packaging-workflows-with-wheel-variants
- https://lwn.net/Articles/1028299/
- Ngôn ngữ chính
- Python
- Star
- 10.6k
- Fork
- 1.5k
- Merge trung bình
- 3 giờ 57 phút
- Pull request đã merge (30 ngày)
- 4
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsCó thể đã có người làm @Belal0066 đã nhận 14 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
abetlen/llama-cpp-python#2371 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2352 · 1 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Docs: consolidate build-from-source and GPU backend guideCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
abetlen/llama-cpp-python#2314 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2211 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyCó thể đã có người làm @Anai-Guo đã nhận 32 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2210 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của abetlen/llama-cpp-python
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 3 ngày
-
Negation with "not" and "no" is ignored during sentiment analysisCó thể đã có người làm @vivek-3728 đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
techcsispit/mess-mood#11 · 1 bình luận ·
-
changelog investigate
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
ramnes/notion-sdk-py#408 ·
-
good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 83/100
btclib-org/btclib-wallet#267 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
good first issue tech-debt
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
knnmelprop/YAADO#111 ·