Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

--vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding

Closed
#767 2 comments 3 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp, python

Research direction

Start with the CPU and GPU vision settings in setup.py and the mtmd argument handling in strata_vision.cpp; compare how --max-tokens is passed and check whether --image-min-tokens can be configured. Also review the CPU encode-time claims in docs/INSTALL.md and docs/TROUBLESHOOTING.md. Done means the CPU path can use a larger image-token budget without requiring that minimum for small images, and the documented CPU encode cost is appropriately qualified or updated.

Written by the indexing model from the issue text.

Description

What

On the AMD backend --vision cpu is the only image path (#304), and VISION["cpu"]["max_tokens"] in
setup.py caps an image at 300 tokens (the GPU path gets 1024). llama.cpp's mtmd prints this at encoder
start:

load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842

strata_vision.cpp exposes --max-tokens but not --image-min-tokens, so that advice cannot be followed
from Strata's config — the 300-token cap is also the minimum mtmd will use here, so the two settings are
tied.

Why it matters

Fine-grained visual questions are where a 300-token grid is weakest: pointing at an object, counting small
items, reading small text inside a large scene, OCR-ish tasks. On this machine the encoder produced 260
tokens for a 640x400 image (grid 20x13) in 2.73 s, so a larger budget is affordable when the image warrants
it — the cost is linear in tokens, and it is paid once per image (results are cached by image hash, so a
conversation sends it once).

Suggestion (not a patch)

Either of these would do, and neither changes the wire format:

  • pass --image-min-tokens through to mtmd alongside the existing --max-tokens, and pick the pair per path
    (e.g. CPU max-tokens 1024 / image-min-tokens 300, so small images stay cheap and large ones get the
    resolution the model wants); or
  • make VISION["cpu"]["max_tokens"] configurable, the way VISION["gpu"] is a table today.

Also: the documented CPU encode cost looks pessimistic

docs/INSTALL.md and docs/TROUBLESHOOTING.md say a picture takes 10-30 s on the CPU encoder. Measured
here (RX 7900 XTX host, engine 0.1.38, Swift 1.5 IQ3_XXS, 8 threads, VISION["cpu"]["max_tokens"]=300,
640x400 PNG): 2.73 s, and 2.74 s on the repeat:

$ printf 'ENC shot.png /tmp/o.sve\nQUIT\n' | engine/strata-vision --mmproj … --model … --max-tokens 300 --threads 8
READY 2560
OK 260 20 13 2727

End to end, a request with an image took 7.0 s wall (2.7 s encode + 321 prompt tokens prefill + ~120 tokens
generated). If 10-30 s was measured on a slower CPU or at the GPU path's 1024 tokens, the docs could say so —
the current figure sets the wrong expectation for anyone deciding whether --vision cpu is usable.

Environment: RX 7900 XTX (gfx1100), Ryzen 7 9800X3D, 60 GB RAM, engine 0.1.38 (99f3dbd), inside
rocm/dev-ubuntu-24.04:10.0.0-full.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.