Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)

Open Beginner friendly
#971 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
72/100
Issue type
Documentation
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp, linux
Domain
documentation

Research direction

Start with docs/UNSLOTH_Q4.md and compare its statement about UD-IQ4_XS with images against the reported 0.1.39 CUDA 12 setup and checks. Verify the setup.py command, VM limitations, and successful image tests in the report; done means the documentation accurately reflects this supported configuration and repeatable path.

Written by the indexing model from the issue text.

Description

docs/UNSLOTH_Q4.md says UD-IQ4_XS with images is "not yet run with images". Here it runs, and it is in daily agent use (Hermes Agent on Windows, over the OpenAI API, driving Unreal Engine work). This report covers what we did to make it run in a VM with a vGPU on an older driver, so others (people and LLMs) can repeat it.

Hardware

GPU NVIDIA L40S-48Q vGPU (VMware ESXi): 48 GB reported, ~42.2 GiB actually allocatable (~5.5 GiB is vGPU/hypervisor overhead)
Driver 550.127.05 (exposes CUDA 12.4)
CPU Intel Xeon Platinum 8580, 32 vCPU, AVX-512
RAM 251 GB
OS Ubuntu 22.04.5 LTS (kernel 6.8), a VM

Running a CUDA 13 project on a CUDA 12 driver

In a vGPU guest, the NVIDIA driver is coupled to the host's vGPU Manager, so it cannot be upgraded from inside the VM. CUDA 13 needs driver 580+, so here CUDA 12.4 is the ceiling.

  • 0.1.38: setup refused drivers below 580 (MIN_DRIVER = 580). We installed the CUDA 12.4 toolkit in a user folder (no system install), built the engine and strata-vision from source (--build, sm_89), and ran setup through a small wrapper. The wrapper put CUDA 12.4 first on PATH / LD_LIBRARY_PATH and lowered MIN_DRIVER to 550 in memory only, without editing the repo. Everything ran on driver 550 via CUDA minor-version compatibility.
  • 0.1.39: this is official now, so the wrapper only sets the toolkit path. --cuda 12 accepts driver 525+ on Linux:
setup.py --setup --family unsloth --model UD-IQ4_XS --context 262144 --kv int8 --vision gpu --cuda 12 --build --no-start --yes

The engine went to engine-cuda12/, so the earlier IQ3_S install stayed intact for rollback. The encoder is the same mmproj-Qwen3.8-Flash-Next-BF16.gguf (SHA-256 checked). Engine log: all 55 GiB of experts in RAM; expert cache 14,672 slots (33.1 GiB of VRAM); 43.4 GB of VRAM in use with images on.

VM / vGPU limitations worth knowing

  • Plan VRAM against ~42.2 GiB, not 48 GB. Tools that size themselves as a fraction of "total" overshoot.
  • The driver cannot be upgraded from the guest: CUDA 13 builds and the ready-made engine are out; the path is CUDA 12 + --build.
  • The GPU may be shared with other services in the same VM: only one GPU workload at a time fits.
  • Some telemetry is missing on vGPU (power, temperature).

Checks (all passed)

  • /health reports images: true at 262144 context.
  • A synthetic image (3 shapes, 3 colours): correct count, shapes and order.
  • A Blender viewport render (a mannequin on a throne): correct count, where the hands are, where the feet are.
  • In the agent: yes/no questions about Unreal viewport captures agreed with pixel measurements. On an all-black capture, the model said it could not see the object instead of inventing it.
  • Also on the same build: tool calls, streaming, and finding a fact at 89.5K tokens of context.

Speed: a small, fair trade

Same machine, images on in both:

IQ3_S (0.1.38) UD-IQ4_XS (0.1.39)
Prompt reading at ~90K context ~3,950 tok/s ~3,300 tok/s
Decode at ~90K context ~93 tok/s ~84 tok/s
Request with one image (incl. answer) — 3.6–4.8 s

About 10% slower decode and ~17% slower prompt reading. In our use, the better answers and visual judgements of the ~4-bit model are worth it. This is a usage impression, not a controlled A/B: a small blind comparison of text and reasoning did not show a measurable difference yet. Some small "filling-in" was seen (4 window panes where there were 2, shading read as body features), so we keep numeric checks as the primary gate and vision as a second opinion.

Tips

  • Keep your client's model name: add "aliases": ["<old model name>"] to the new strata-*.json, so a client configured with that name needs no change.
  • Carry your edits over: setup writes a new strata-*.json per model. Copy your own edits (api_key, host, sampling, reasoning_budget_tokens) from the old one.
  • The <think> fix: see #804 / #970 for the fix for tool calls written inside the thinking (seen with this model family in agent use).

Next, I will test even larger quantizations on this machine and report back.

🤖 Generated with Claude Code

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.