[Feature]: AMD Strix Halo Vulkan GPU Acceleration and Ollama Engine Configuration
Les mainteneurs répondent en général sous 5 jours
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 25/100
Piste de recherche
Start with how the router launches its managed Ollama engine on port 11435 and which environment it passes to that process; the reporter says no supported setting exists for this. Check that OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1 reach the child process, then confirm with size_vram that the model loads on the Radeon 8060S. Done means a persisted, validated setting that enables Vulkan without changing cluster routing or the external API; the maintainers still need to decide the design.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Area
Engine or model management
User problem
Is AMD Strix Halo (Ryzen AI Max+ 395 / Radeon 8060S) running Ubuntu Linux an intended PAIR inference platform? We have confirmed that PAIR 0.1.1 operates successfully on this system, and its bundled Ollama 0.34.1 can run Qwen 3.6 27B with Vulkan acceleration when OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1 are supplied. However, PAIR's managed Ollama starts without those settings and runs inference on the CPU. Is there a supported way to enable Vulkan for the managed engine?
Desired outcome
Provide a supported method for enabling GPU acceleration in PAIR-managed Ollama on compatible AMD hardware, including AMD Strix Halo (Ryzen AI Max+ 395 / Radeon 8060S).
Ideally, PAIR would allow users to configure the inference backend or supply engine-specific environment variables, such as OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1, through the application settings.
These settings should persist across engine restarts and application updates without requiring modifications to PAIR's internal files.
The expected result is that supported AMD GPUs can execute model inference through PAIR with GPU acceleration, while preserving existing cluster routing and model management functionality.
We would also appreciate clarification on whether AMD Strix Halo is an officially supported or planned GPU inference platform.
Alternatives considered
We successfully tested PAIR's bundled Ollama 0.34.1 independently using a temporary server on port 11436, with OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1.
The test detected the Radeon 8060S through Vulkan and successfully ran Qwen 3.6 27B with the full model allocation reported on the GPU.
In contrast, PAIR's managed Ollama instance on port 11435 runs the same model on the CPU, with size_vram reported as 0.
We examined the available PAIR configuration files, engine manager command-line options, and desktop application settings. We did not identify a supported mechanism for configuring GPU acceleration or passing environment variables to the managed Ollama engine.
We considered maintaining a separate GPU-enabled Ollama instance or modifying PAIR's startup environment, but neither approach has been validated for our existing multi-node PAIR cluster. We prefer a supported configuration method.
Compatibility and security implications
The requested functionality should preserve PAIR's existing proxy, networking, cluster pairing, authentication, and model-routing behavior.
GPU backend selection and engine environment variables should apply only to the locally managed inference engine, without changing the cluster's external API endpoints or exposing additional network services.
Existing configurations should continue to work without modification. GPU acceleration should remain optional, with current defaults preserved for systems where the requested backend is unavailable.
Any engine settings should persist across application restarts and updates. Ideally, the interface would validate supported options and avoid exposing arbitrary environment variables that could affect security-sensitive engine behavior.
No changes to model data handling, cluster identity, or authentication are requested.
Validation approach
Suggested validation on an AMD Ryzen AI Max+ 395 / Radeon 8060S system running Ubuntu Linux:
- Enable Vulkan acceleration for PAIR-managed Ollama through the supported configuration interface.
- Restart the managed engine and verify that the configuration persists.
- Confirm that Ollama discovers the Radeon 8060S using the Vulkan backend.
- Load Qwen 3.6 27B (Q4_K_M) through PAIR and execute a short inference request.
- Verify that Ollama reports nonzero size_vram and that AMDGPU memory usage increases during model loading.
- Confirm that PAIR's proxy, model discovery, cluster routing, and model ejection continue functioning normally.
- Restart PAIR and verify that GPU acceleration remains enabled without manual intervention.
For reference, we independently tested PAIR's bundled Ollama 0.34.1 using OLLAMA_VULKAN=1 and OLLAMA_IGPU_ENABLE=1.
The standalone test successfully detected the Radeon 8060S through Vulkan, reported approximately 94.7 GiB of GPU-addressable memory, and ran Qwen 3.6 27B with approximately 17.15 GB allocated to the GPU.
The same model running through PAIR's default managed engine reported size_vram=0.
Confirmations
- I searched existing issues for duplicates.
- I agree to follow the Code of Conduct.
- Langage dominant
- Go
- Étoiles
- 1.6k
- Forks
- 266
- Merge moyen
- 3 j 15 h
- PR mergées (30 j)
- 23
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de NVIDIA/Personal-AI-Router
-
enhancement
Difficulté 4/5 Plus d'une semaine Accessibilité débutants 12/100
NVIDIA/Personal-AI-Router#162 ·
Les mainteneurs répondent en général sous 5 jours
-
[Feature]: [Ollama] Alert the user about model pull failuresPeut-être pris @ckelseynv l’a pris il y a 1 jour. Ouverteenhancement
Difficulté 3/5 1-2 jours Accessibilité débutants 56/100
NVIDIA/Personal-AI-Router#152 · 1 commentaire · 1 personne assignée ·
Les mainteneurs répondent en général sous 5 jours
-
[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxyPeut-être pris @ilaigold l’a pris il y a 2 jours. Ouverte
Difficulté 3/5 Une demi-journée Accessibilité débutants 84/100
NVIDIA/Personal-AI-Router#146 ·
Les mainteneurs répondent en général sous 5 jours
-
Make model download cancellation asynchronousPeut-être pris @ckelseynv l’a pris il y a 18 jours. Ouverteenhancement
NVIDIA/Personal-AI-Router#115 · 1 personne assignée ·
Les mainteneurs répondent en général sous 5 jours
-
Derive the model download and cancellation timeouts from measurementPeut-être pris @ckelseynv l’a pris il y a 18 jours. Ouverteenhancement
NVIDIA/Personal-AI-Router#114 · 1 personne assignée ·
Les mainteneurs répondent en général sous 5 jours
Toutes les issues de NVIDIA/Personal-AI-Router
Issues similaires
-
bug go
Difficulté 2/5 1-3 heures Accessibilité débutants 82/100
genkit-ai/genkit#6761 · 1 commentaire ·
Les mainteneurs répondent en général sous 2 jours
-
bug(backend): `make test-update` in backend/src/v2 fails because the --update flag was removedOuverteready
Difficulté 2/5 1-3 heures Accessibilité débutants 92/100
kubeflow/pipelines#14784 · 1 commentaire ·
Les mainteneurs répondent en général sous 2 jours
-
bug frontend good first issue
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
Les mainteneurs répondent en général sous 1 jour
-
trust: update-propagation-directive requires developer mode while add and remove do notPeut-être pris @bhuvan-somisetty l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
Les mainteneurs répondent en général sous 2 jours
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 82/100
oalders/clodhopper#133 ·