Benchmark local models and GPUs for the reports and the behaviour check
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 68/100
- Tipo di issue
- Documentazione
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Ambito
- performance
Direzione di ricerca
Follow the issue steps literally: check out PR #110, run llama-bench for raw speed, then start llama-server with --jinja and scripts/gemini-shim.py with --thinking off. Save the given local-bench.sh and run it twice; it exercises cargo test --lib a_local_model_writes_a_report and cargo test --test interview_behavior. Done means posting a comment filling the results template (hardware fields plus report timings from the shim log and behaviour-check pass counts), attaching the full log on failure.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
#110 lets the report, the pause review, the phase judge and the interviewer behaviour check run on a model you host yourself, through scripts/gemini-shim.py in front of llama.cpp's llama-server. So far it has been measured on one GPU with one model. This issue collects results from other models and other hardware, so that we can answer two questions:
- Which models follow the interview rules? The behaviour check fails a model that names the source problem, volunteers a limit, answers a hint request without calling
log_hint, or serves a hint past its rung. - Which hardware is fast enough? The local deadlines in #110 are fixed: 45 s for each report call, 250 s for the whole report, and 12 s for each pause review or phase judgment. A slower GPU, or a model that writes longer answers, will hit them.
The live interviewer is not covered. In #110 it still talks to Gemini Live, so the steps below need neither a Gemini key nor LiveKit.
What you need
- A GPU that llama.cpp supports (CUDA, ROCm, Vulkan or Metal), or a CPU if you are patient. gemma-4-12b Q4_K_M with a 32k context takes about 8.2 GB of VRAM.
- Rust 1.98 or newer and Python 3.9 or newer. On Linux x86_64,
make builddownloads the Clang that WebRTC needs; on other hosts, see the README's "Dependencies by lifecycle". - About 20 GB of disk: the model, a llama.cpp build and CodeTrial's
target/.
Steps
1. Check out #110 and build it once
git clone https://github.com/sysprog21/codetrial && cd codetrial
gh pr checkout 110 # or: git fetch origin pull/110/head:pr-110 && git checkout pr-110
make build
2. Build llama.cpp and measure the raw speed
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON # -DGGML_HIP=ON, -DGGML_VULKAN=ON, or nothing on macOS for Metal
cmake --build build -j --config Release
build/bin/llama-bench -m MODEL.gguf -p 512,4096 -n 256 -fa 1
Keep the llama-bench table for the results.
3. Start the model and the shim, each in its own terminal:
build/bin/llama-server -m MODEL.gguf --host 127.0.0.1 --port 8080 \
-ngl 99 -c 32768 -fa on -ctk q8_0 -ctv q8_0 --jinja
scripts/gemini-shim.py --listen 127.0.0.1:8090 --llama http://127.0.0.1:8080 --thinking off
--jinja is required, or the tool calls do not work. Use --thinking off for models that think by default (Gemma 4, Qwen3): with thinking on, gemma-4-12b sometimes repeated itself until the output limit and took about ten times as long. If the model does not fit, lower -c or -ngl and note what you used.
4. Run the bench from the CodeTrial checkout. Save this as local-bench.sh:
#!/bin/sh
# Report x3, then the behaviour check (which includes the phase judge).
set -u
export CODETRIAL_GEMINI_REST_BASE=${SHIM:-http://127.0.0.1:8090}
export BEHAVIOR_CANDIDATE_BASE=${LLAMA:-http://127.0.0.1:8080}
export GOOGLE_API_KEY=local
export CXX=${CXX:-$PWD/target/clang/bin/clang++}
log=local-bench-$(date +%Y%m%d-%H%M).log
for i in 1 2 3; do
echo "== report $i"
cargo test -q --lib a_local_model_writes_a_report -- --ignored --nocapture
done >"$log" 2>&1
echo "== behaviour" >>"$log"
cargo test -q --test interview_behavior -- --ignored --nocapture >>"$log" 2>&1
grep -E '^== |^elapsed|--- FAILED|^test result' "$log"
grep -E '^[a-z0-9-]+(/[a-z-]+| \([A-Za-z]+\)): ' "$log"
grep -A1 'panicked at' "$log" | grep -vE 'panicked at|^--$|behaviour failures'
echo "full log: $log"
sh local-bench.sh
sh local-bench.sh
Please run it twice: results vary between runs even though the calls are seeded (see the reference below). On an RTX 5070 Ti one run takes about 5 minutes. The report test uses the Two Sum golden prompt by default, and REPORT_PROBLEM changes the problem. The shim's own log shows each call's tokens and time; please include the report calls (... in, ... out, STOP, ...s).
Results template
Post a comment with this filled in, and attach the full log if anything failed. These commands print most of the hardware fields:
# Linux
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv # or rocm-smi / vulkaninfo --summary
lscpu | grep "Model name"; free -g | grep Mem; cat /etc/os-release | grep PRETTY
# macOS
system_profiler SPHardwareDataType SPDisplaysDataType | grep -E "Chip|Memory|Total Number of Cores"; sw_vers
GPU (model, VRAM):
CPU:
RAM:
OS:
GPU backend and driver: (e.g. CUDA 12.9 / driver 580, ROCm 6.4, Vulkan, Metal)
llama.cpp commit:
Model file: (e.g. gemma-4-12b-it-Q4_K_M.gguf)
llama-server flags:
llama-bench pp512 / pp4096 / tg256:
VRAM in use while serving:
For each of the two runs:
Report x3: valid? / seconds each / calls each (from the shim log)
Behaviour check: tests passed of 5 (the phase judge is one of them),
and the failure lines the script prints
Anything else you saw:
Reference result
RTX 5070 Ti 16 GB, i7-14700, 64 GB RAM, Ubuntu 24.04, driver 580.173, llama.cpp bd43117, gemma-4-12b-it Q4_K_M with the flags above and --thinking off:
- Throughput from the server log: prompt about 3,400 tokens/s, generation about 77 tokens/s.
| Run 1 | Run 2 | |
|---|---|---|
| Report x3 | all valid, 44.5 s each, two calls each: the first answer (3,267 tokens in, 1,846 out, 24 s) was refused by validation and the repair (5,431 in, 1,530 out, 20.5 s) was accepted | all valid, 15 to 16 s each, one call each |
| Behaviour check | 2 of 5 | 3 of 5 |
| Phase judge test | pass | pass |
Hint order (live_interviewer_poses_the_variant_and_serves_hints_in_order) |
fail: a second hint request answered without log_hint, on 3sum and on two-sum with the whiteboard |
pass |
| Uncertain speech | fail: wrong indices not challenged against the input | fail, the same way |
| Played candidates | fail, 5 lines | fail, 8 lines: hint requests answered without log_hint, and limits volunteered (3sum 100,000, coin-change 10,000, two-sum 10^9) |
The failures are model behaviour, not timeouts. Every report call stayed inside the 45 s limit, the longest at 24 s.
In a five-minute interview by hand on the same machine, with the interviewer also local, a report took one call and 17 s.
Worth trying
- Models: Qwen3-14B, Qwen3.5-9B, gpt-oss-20b, Mistral Small 3.x, Llama 3.x 8B, and larger quantizations of gemma-4-12b. Earlier trials on the interviewer side here: Qwen3.5-9B named the intended data structure unasked, and Qwen3-14B followed the hint rules best but did not fit fully in 16 GB beside the speech models.
- Hardware: 8 GB and 12 GB NVIDIA cards, AMD through ROCm or Vulkan, and Apple Silicon. On an M1 Pro, my estimate is that a report call takes 90 s or more, which would hit the 45 s limit; a measurement would settle it. If it does, the deadlines should probably be configurable for local setups, and the results here are what would size them.
Related: #105 (running CodeTrial on a local model), #110.
- Lingua principale
- Rust
- Stelle
- 144
- Fork
- 41
- Merge medio
- 1g 7h
- PR unite (30g)
- 94
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di sysprog21/codetrial
-
Highlight the active line in the code editorForse già presa @ArthurArthurArthur0817 l’ha presa 2 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
sysprog21/codetrial#247 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Support configurable and randomized interviewer voices and accentsForse già presa @MorganHo001 l’ha presa 1 giorno fa. Aperta
sysprog21/codetrial#261 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
sysprog21/codetrial#258 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Add an example drawing area to coding interviewsForse già presa @Chuyutseng l’ha presa 1 giorno fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
sysprog21/codetrial#255 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Add indentation guides to the code editorForse già presa @sidney-Hung l’ha presa 2 giorni fa. Aperta
sysprog21/codetrial#254 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di sysprog21/codetrial
Issue simili
-
enhancement user-priority/P3
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 75/100
element-hq/lk-jwt-service#248 ·
I maintainer di solito rispondono entro 1 giorno
-
agent:triaged bug bughunt pm:pipenv priority:p1
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
SocketDev/socket-patch#1219 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
pact-foundation/pact-cli#154 ·
I maintainer di solito rispondono entro 3 giorni