Add support for experimental wheel variants (i.e., wheelnext)
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 45/100
- Issue 類型
- 功能
- 描述清晰度
- 基本清楚
- 活躍度
- 停滯
- 技術堆疊
- python
- 領域
- build-system, tooling
研究方向
這個 issue 是關於修改 wheel 的建置與發佈流程。首先檢查專案的建置腳本,可能位於 setup.py 或 pyproject.toml 中,以及 CI/CD 工作流程。研究 WheelNext 規範,以及 PyTorch 等專案如何實作變體中繼資料。目標是產生具有正確中繼資料的特定後端 wheel(CUDA、ROCm、Metal),同時確保 CPU wheel 仍作為 fallback。測試將包括在本機建置 wheel 並驗證中繼資料。
由索引模型根據 Issue 內容生成。
描述
Is your feature request related to a problem? Please describe.
Today, installing llama-cpp-python on machines with different GPU backends (CUDA, ROCm, Metal, etc.) requires separate package names, custom extra indexes, or installer-level logic to select the correct wheel. This creates friction for downstream tooling (CLIs, orchestrators, and packaging systems) that want to provide a “just works” experience, especially when users don’t know which backend they need. Even a simple developer-driven install might require picking precisely the correct wheel.
Describe the solution you'd like
Add support for WheelNext-compatible experimental wheel variants when building and publishing wheels.
This would allow llama-cpp-python to produce a single package version that provides multiple backend-aware binary wheels, each annotated with variant metadata (e.g., GPU type, CUDA version, ROCm version).
Installers that understand the WheelNext spec (now used experimentally by PyTorch, uv, and others) can automatically select the correct backend wheel based on the system’s hardware/software configuration without a need for custom index URLs, separate packages, or manual backend flags.
Key pieces:
- Generate wheels with variant metadata following the experimental WheelNext (wheel variants) conventions.
- Publish per-backend wheels using the standardized naming + metadata fields.
- Ensure that CPU-only wheels remain available as fallback.
This would significantly simplify installation for all users and remove backend-selection logic from downstream tools. Wheel variants are fully backward-compatible so existing workflows won't be disrupted.
Describe alternatives you've considered
- Separate package names per backend (e.g., llama-cpp-python-cuda): fragments packaging and forces manual selection.
- Extras for backend variants (pip install llama-cpp-python[cuda]): still requires external detection and doesn’t integrate with hardware-aware installer selection.
- Custom index URLs for backend wheels: brittle and requires orchestration logic outside Python packaging.
- CLI-backed installation routing (what many downstream projects do currently): it’s reinventing the wheel and provides an inconsistent experience for end users.
All of these solutions put the burden on downstream tooling rather than on standardized wheel metadata.
Additional context
- https://wheelnext.dev
- https://pytorch.org/blog/pytorch-wheel-variants/
- https://astral.sh/blog/wheel-variants
- https://labs.quansight.org/blog/python-wheels-from-tags-to-variants
- https://developer.nvidia.com/blog/streamline-cuda-accelerated-python-install-and-packaging-workflows-with-wheel-variants
- https://lwn.net/Articles/1028299/
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.5k
- 平均合併
- 3 小時 57 分鐘
- 30 天內合併 PR
- 4
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
abetlen/llama-cpp-python 的其他 Issue
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently drops可能已有人在做 @Belal0066 於 16 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
維護者通常 1 天內回覆
-
uv add llama-cpp-python wheels fails for versions above 0.3.30可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 則留言 · 2 個 reaction ·
維護者通常 1 天內回覆
-
Docs: consolidate build-from-source and GPU backend guide可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
維護者通常 1 天內回覆
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_array可能重新可做 @lxcxjxhx 於 93 天前認領,目前沒有進行中的 PR。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 則留言 ·
維護者通常 1 天內回覆
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusingly可能已有人在做 @Anai-Guo 於 34 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
維護者通常 1 天內回覆
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
bug
難度 2/5 1-3 小時 新手友好度 85/100
Deepak3699/Ai_Mentor#244 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 82/100
btclib-org/btclib-node#1880 ·
維護者通常 1 天內回覆
-
CONTRIBUTING.md: say how ticketless bug fixes and feature PRs are handled可能已有人在做 @khuisman 今天認領。 未關閉v0.9.2
難度 1/5 1 小時以內 新手友好度 84/100
khuisman/mcp-gee-sweet#941 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 68/100
pyjanitor-devs/pyjanitor#1758 ·
維護者通常 1 天內回覆
-
bug ready for review
難度 2/5 1-3 小時 新手友好度 86/100
odysseus-dev/odysseus#6641 ·
維護者通常 1 天內回覆