Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 72/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp, linux
- Domain
- documentation
Research direction
Start with docs/UNSLOTH_Q4.md and compare its statement about UD-IQ4_XS with images against the reported 0.1.39 CUDA 12 setup and checks. Verify the setup.py command, VM limitations, and successful image tests in the report; done means the documentation accurately reflects this supported configuration and repeatable path.
Written by the indexing model from the issue text.
Description
docs/UNSLOTH_Q4.md says UD-IQ4_XS with images is "not yet run with images". Here it runs, and it is in daily agent use (Hermes Agent on Windows, over the OpenAI API, driving Unreal Engine work). This report covers what we did to make it run in a VM with a vGPU on an older driver, so others (people and LLMs) can repeat it.
Hardware
| GPU | NVIDIA L40S-48Q vGPU (VMware ESXi): 48 GB reported, ~42.2 GiB actually allocatable (~5.5 GiB is vGPU/hypervisor overhead) |
| Driver | 550.127.05 (exposes CUDA 12.4) |
| CPU | Intel Xeon Platinum 8580, 32 vCPU, AVX-512 |
| RAM | 251 GB |
| OS | Ubuntu 22.04.5 LTS (kernel 6.8), a VM |
Running a CUDA 13 project on a CUDA 12 driver
In a vGPU guest, the NVIDIA driver is coupled to the host's vGPU Manager, so it cannot be upgraded from inside the VM. CUDA 13 needs driver 580+, so here CUDA 12.4 is the ceiling.
- 0.1.38: setup refused drivers below 580 (
MIN_DRIVER = 580). We installed the CUDA 12.4 toolkit in a user folder (no system install), built the engine andstrata-visionfrom source (--build,sm_89), and ran setup through a small wrapper. The wrapper put CUDA 12.4 first onPATH/LD_LIBRARY_PATHand loweredMIN_DRIVERto 550 in memory only, without editing the repo. Everything ran on driver 550 via CUDA minor-version compatibility. - 0.1.39: this is official now, so the wrapper only sets the toolkit path.
--cuda 12accepts driver 525+ on Linux:
setup.py --setup --family unsloth --model UD-IQ4_XS --context 262144 --kv int8 --vision gpu --cuda 12 --build --no-start --yes
The engine went to engine-cuda12/, so the earlier IQ3_S install stayed intact for rollback. The encoder is the same mmproj-Qwen3.8-Flash-Next-BF16.gguf (SHA-256 checked). Engine log: all 55 GiB of experts in RAM; expert cache 14,672 slots (33.1 GiB of VRAM); 43.4 GB of VRAM in use with images on.
VM / vGPU limitations worth knowing
- Plan VRAM against ~42.2 GiB, not 48 GB. Tools that size themselves as a fraction of "total" overshoot.
- The driver cannot be upgraded from the guest: CUDA 13 builds and the ready-made engine are out; the path is CUDA 12 +
--build. - The GPU may be shared with other services in the same VM: only one GPU workload at a time fits.
- Some telemetry is missing on vGPU (power, temperature).
Checks (all passed)
/healthreportsimages: trueat 262144 context.- A synthetic image (3 shapes, 3 colours): correct count, shapes and order.
- A Blender viewport render (a mannequin on a throne): correct count, where the hands are, where the feet are.
- In the agent: yes/no questions about Unreal viewport captures agreed with pixel measurements. On an all-black capture, the model said it could not see the object instead of inventing it.
- Also on the same build: tool calls, streaming, and finding a fact at 89.5K tokens of context.
Speed: a small, fair trade
Same machine, images on in both:
| IQ3_S (0.1.38) | UD-IQ4_XS (0.1.39) | |
|---|---|---|
| Prompt reading at ~90K context | ~3,950 tok/s | ~3,300 tok/s |
| Decode at ~90K context | ~93 tok/s | ~84 tok/s |
| Request with one image (incl. answer) | — | 3.6–4.8 s |
About 10% slower decode and ~17% slower prompt reading. In our use, the better answers and visual judgements of the ~4-bit model are worth it. This is a usage impression, not a controlled A/B: a small blind comparison of text and reasoning did not show a measurable difference yet. Some small "filling-in" was seen (4 window panes where there were 2, shading read as body features), so we keep numeric checks as the primary gate and vision as a second opinion.
Tips
- Keep your client's model name: add
"aliases": ["<old model name>"]to the newstrata-*.json, so a client configured with that name needs no change. - Carry your edits over: setup writes a new
strata-*.jsonper model. Copy your own edits (api_key,host,sampling,reasoning_budget_tokens) from the old one. - The
<think>fix: see #804 / #970 for the fix for tool calls written inside the thinking (seen with this model family in agent use).
Next, I will test even larger quantizations on this machine and report back.
🤖 Generated with Claude Code
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
cataclysmbn/Cataclysm-BN#10516 ·
Maintainers usually reply within 1 day
-
area:runtime good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
WATonomous/wato_f1tenth#39 ·
-
[APP BUG]: Sorting by name after searching can bring up irrelevant resultsPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
duckdb/duckdb-excel#104 ·