--vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- backend, documentation
Research direction
Start with the CPU and GPU vision settings in setup.py and the mtmd argument handling in strata_vision.cpp; compare how --max-tokens is passed and check whether --image-min-tokens can be configured. Also review the CPU encode-time claims in docs/INSTALL.md and docs/TROUBLESHOOTING.md. Done means the CPU path can use a larger image-token budget without requiring that minimum for small images, and the documented CPU encode cost is appropriately qualified or updated.
Written by the indexing model from the issue text.
Description
What
On the AMD backend --vision cpu is the only image path (#304), and VISION["cpu"]["max_tokens"] in
setup.py caps an image at 300 tokens (the GPU path gets 1024). llama.cpp's mtmd prints this at encoder
start:
load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842
strata_vision.cpp exposes --max-tokens but not --image-min-tokens, so that advice cannot be followed
from Strata's config — the 300-token cap is also the minimum mtmd will use here, so the two settings are
tied.
Why it matters
Fine-grained visual questions are where a 300-token grid is weakest: pointing at an object, counting small
items, reading small text inside a large scene, OCR-ish tasks. On this machine the encoder produced 260
tokens for a 640x400 image (grid 20x13) in 2.73 s, so a larger budget is affordable when the image warrants
it — the cost is linear in tokens, and it is paid once per image (results are cached by image hash, so a
conversation sends it once).
Suggestion (not a patch)
Either of these would do, and neither changes the wire format:
- pass
--image-min-tokensthrough to mtmd alongside the existing--max-tokens, and pick the pair per path
(e.g. CPUmax-tokens 1024/image-min-tokens 300, so small images stay cheap and large ones get the
resolution the model wants); or - make
VISION["cpu"]["max_tokens"]configurable, the wayVISION["gpu"]is a table today.
Also: the documented CPU encode cost looks pessimistic
docs/INSTALL.md and docs/TROUBLESHOOTING.md say a picture takes 10-30 s on the CPU encoder. Measured
here (RX 7900 XTX host, engine 0.1.38, Swift 1.5 IQ3_XXS, 8 threads, VISION["cpu"]["max_tokens"]=300,
640x400 PNG): 2.73 s, and 2.74 s on the repeat:
$ printf 'ENC shot.png /tmp/o.sve\nQUIT\n' | engine/strata-vision --mmproj … --model … --max-tokens 300 --threads 8
READY 2560
OK 260 20 13 2727
End to end, a request with an image took 7.0 s wall (2.7 s encode + 321 prompt tokens prefill + ~120 tokens
generated). If 10-30 s was measured on a slower CPU or at the GPU path's 1024 tokens, the docs could say so —
the current figure sets the wrong expectation for anyone deciding whether --vision cpu is usable.
Environment: RX 7900 XTX (gfx1100), Ryzen 7 9800X3D, 60 GB RAM, engine 0.1.38 (99f3dbd), inside
rocm/dev-ubuntu-24.04:10.0.0-full.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
EsotericSoftware/spine-runtimes#3186 ·
-
An empty line splits a signature where an ordinary comment is right above an argument's HaddockOpen
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
mrkkrp/tilia#213 · 1 comment ·
Maintainers usually reply within 1 day
-
Round video messages start gray and blocky with libx264: encoder is configured for 1,000,000 fpsOpen
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
telegramdesktop/tdesktop#31422 ·
Maintainers usually reply within 9 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 5 days