Add support for experimental wheel variants (i.e., wheelnext)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 45/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Estancado
- Stack tecnológico
- python
- Área
- build-system, tooling
Línea de trabajo
El issue trata sobre modificar el proceso de compilación y publicación de wheels. Empieza examinando los scripts de build del proyecto, probablemente en setup.py o pyproject.toml, y los workflows de CI/CD. Investiga la especificación de WheelNext y cómo proyectos como PyTorch implementan los metadatos de variantes. El objetivo es producir wheels específicos del backend (CUDA, ROCm, Metal) con los metadatos correctos, asegurando que los wheels de CPU sigan disponibles como fallback. Las pruebas consistirán en compilar wheels localmente y verificar los metadatos.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Is your feature request related to a problem? Please describe.
Today, installing llama-cpp-python on machines with different GPU backends (CUDA, ROCm, Metal, etc.) requires separate package names, custom extra indexes, or installer-level logic to select the correct wheel. This creates friction for downstream tooling (CLIs, orchestrators, and packaging systems) that want to provide a “just works” experience, especially when users don’t know which backend they need. Even a simple developer-driven install might require picking precisely the correct wheel.
Describe the solution you'd like
Add support for WheelNext-compatible experimental wheel variants when building and publishing wheels.
This would allow llama-cpp-python to produce a single package version that provides multiple backend-aware binary wheels, each annotated with variant metadata (e.g., GPU type, CUDA version, ROCm version).
Installers that understand the WheelNext spec (now used experimentally by PyTorch, uv, and others) can automatically select the correct backend wheel based on the system’s hardware/software configuration without a need for custom index URLs, separate packages, or manual backend flags.
Key pieces:
- Generate wheels with variant metadata following the experimental WheelNext (wheel variants) conventions.
- Publish per-backend wheels using the standardized naming + metadata fields.
- Ensure that CPU-only wheels remain available as fallback.
This would significantly simplify installation for all users and remove backend-selection logic from downstream tools. Wheel variants are fully backward-compatible so existing workflows won't be disrupted.
Describe alternatives you've considered
- Separate package names per backend (e.g., llama-cpp-python-cuda): fragments packaging and forces manual selection.
- Extras for backend variants (pip install llama-cpp-python[cuda]): still requires external detection and doesn’t integrate with hardware-aware installer selection.
- Custom index URLs for backend wheels: brittle and requires orchestration logic outside Python packaging.
- CLI-backed installation routing (what many downstream projects do currently): it’s reinventing the wheel and provides an inconsistent experience for end users.
All of these solutions put the burden on downstream tooling rather than on standardized wheel metadata.
Additional context
- https://wheelnext.dev
- https://pytorch.org/blog/pytorch-wheel-variants/
- https://astral.sh/blog/wheel-variants
- https://labs.quansight.org/blog/python-wheels-from-tags-to-variants
- https://developer.nvidia.com/blog/streamline-cuda-accelerated-python-install-and-packaging-workflows-with-wheel-variants
- https://lwn.net/Articles/1028299/
- Lenguaje dominante
- Python
- Estrellas
- 10.6k
- Forks
- 1.5k
- Merge medio
- 3 h 57 min
- PR fusionados (30 d)
- 4
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsPosiblemente ocupada @Belal0066 la tomó hace 13 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
abetlen/llama-cpp-python#2371 ·
Los mantenedores suelen responder en 1 día
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Posiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2352 · 1 comentario · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
Docs: consolidate build-from-source and GPU backend guidePosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
abetlen/llama-cpp-python#2314 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2211 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyPosiblemente ocupada @Anai-Guo la tomó hace 31 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
abetlen/llama-cpp-python#2210 ·
Los mantenedores suelen responder en 1 día
Todos los issues de abetlen/llama-cpp-python
Issues similares
-
needs-human needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
gke-labs/kube-agents#2400 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedPosiblemente ocupada @sylvesterkaczmarek la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
google-deepmind/bsuite#56 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
LearningCircuit/local-deep-research#7206 ·
Los mantenedores suelen responder en 1 día
-
[TASK] Document technology stackAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Los mantenedores suelen responder en 1 día