Benchmark local models and GPUs for the reports and the behaviour check
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 68/100
- issue の種類
- ドキュメント
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
- 領域
- performance
調査の方向性
Follow the issue steps literally: check out PR #110, run llama-bench for raw speed, then start llama-server with --jinja and scripts/gemini-shim.py with --thinking off. Save the given local-bench.sh and run it twice; it exercises cargo test --lib a_local_model_writes_a_report and cargo test --test interview_behavior. Done means posting a comment filling the results template (hardware fields plus report timings from the shim log and behaviour-check pass counts), attaching the full log on failure.
索引モデルが issue の本文から書いたものです。
説明
#110 lets the report, the pause review, the phase judge and the interviewer behaviour check run on a model you host yourself, through scripts/gemini-shim.py in front of llama.cpp's llama-server. So far it has been measured on one GPU with one model. This issue collects results from other models and other hardware, so that we can answer two questions:
- Which models follow the interview rules? The behaviour check fails a model that names the source problem, volunteers a limit, answers a hint request without calling
log_hint, or serves a hint past its rung. - Which hardware is fast enough? The local deadlines in #110 are fixed: 45 s for each report call, 250 s for the whole report, and 12 s for each pause review or phase judgment. A slower GPU, or a model that writes longer answers, will hit them.
The live interviewer is not covered. In #110 it still talks to Gemini Live, so the steps below need neither a Gemini key nor LiveKit.
What you need
- A GPU that llama.cpp supports (CUDA, ROCm, Vulkan or Metal), or a CPU if you are patient. gemma-4-12b Q4_K_M with a 32k context takes about 8.2 GB of VRAM.
- Rust 1.98 or newer and Python 3.9 or newer. On Linux x86_64,
make builddownloads the Clang that WebRTC needs; on other hosts, see the README's "Dependencies by lifecycle". - About 20 GB of disk: the model, a llama.cpp build and CodeTrial's
target/.
Steps
1. Check out #110 and build it once
git clone https://github.com/sysprog21/codetrial && cd codetrial
gh pr checkout 110 # or: git fetch origin pull/110/head:pr-110 && git checkout pr-110
make build
2. Build llama.cpp and measure the raw speed
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON # -DGGML_HIP=ON, -DGGML_VULKAN=ON, or nothing on macOS for Metal
cmake --build build -j --config Release
build/bin/llama-bench -m MODEL.gguf -p 512,4096 -n 256 -fa 1
Keep the llama-bench table for the results.
3. Start the model and the shim, each in its own terminal:
build/bin/llama-server -m MODEL.gguf --host 127.0.0.1 --port 8080 \
-ngl 99 -c 32768 -fa on -ctk q8_0 -ctv q8_0 --jinja
scripts/gemini-shim.py --listen 127.0.0.1:8090 --llama http://127.0.0.1:8080 --thinking off
--jinja is required, or the tool calls do not work. Use --thinking off for models that think by default (Gemma 4, Qwen3): with thinking on, gemma-4-12b sometimes repeated itself until the output limit and took about ten times as long. If the model does not fit, lower -c or -ngl and note what you used.
4. Run the bench from the CodeTrial checkout. Save this as local-bench.sh:
#!/bin/sh
# Report x3, then the behaviour check (which includes the phase judge).
set -u
export CODETRIAL_GEMINI_REST_BASE=${SHIM:-http://127.0.0.1:8090}
export BEHAVIOR_CANDIDATE_BASE=${LLAMA:-http://127.0.0.1:8080}
export GOOGLE_API_KEY=local
export CXX=${CXX:-$PWD/target/clang/bin/clang++}
log=local-bench-$(date +%Y%m%d-%H%M).log
for i in 1 2 3; do
echo "== report $i"
cargo test -q --lib a_local_model_writes_a_report -- --ignored --nocapture
done >"$log" 2>&1
echo "== behaviour" >>"$log"
cargo test -q --test interview_behavior -- --ignored --nocapture >>"$log" 2>&1
grep -E '^== |^elapsed|--- FAILED|^test result' "$log"
grep -E '^[a-z0-9-]+(/[a-z-]+| \([A-Za-z]+\)): ' "$log"
grep -A1 'panicked at' "$log" | grep -vE 'panicked at|^--$|behaviour failures'
echo "full log: $log"
sh local-bench.sh
sh local-bench.sh
Please run it twice: results vary between runs even though the calls are seeded (see the reference below). On an RTX 5070 Ti one run takes about 5 minutes. The report test uses the Two Sum golden prompt by default, and REPORT_PROBLEM changes the problem. The shim's own log shows each call's tokens and time; please include the report calls (... in, ... out, STOP, ...s).
Results template
Post a comment with this filled in, and attach the full log if anything failed. These commands print most of the hardware fields:
# Linux
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv # or rocm-smi / vulkaninfo --summary
lscpu | grep "Model name"; free -g | grep Mem; cat /etc/os-release | grep PRETTY
# macOS
system_profiler SPHardwareDataType SPDisplaysDataType | grep -E "Chip|Memory|Total Number of Cores"; sw_vers
GPU (model, VRAM):
CPU:
RAM:
OS:
GPU backend and driver: (e.g. CUDA 12.9 / driver 580, ROCm 6.4, Vulkan, Metal)
llama.cpp commit:
Model file: (e.g. gemma-4-12b-it-Q4_K_M.gguf)
llama-server flags:
llama-bench pp512 / pp4096 / tg256:
VRAM in use while serving:
For each of the two runs:
Report x3: valid? / seconds each / calls each (from the shim log)
Behaviour check: tests passed of 5 (the phase judge is one of them),
and the failure lines the script prints
Anything else you saw:
Reference result
RTX 5070 Ti 16 GB, i7-14700, 64 GB RAM, Ubuntu 24.04, driver 580.173, llama.cpp bd43117, gemma-4-12b-it Q4_K_M with the flags above and --thinking off:
- Throughput from the server log: prompt about 3,400 tokens/s, generation about 77 tokens/s.
| Run 1 | Run 2 | |
|---|---|---|
| Report x3 | all valid, 44.5 s each, two calls each: the first answer (3,267 tokens in, 1,846 out, 24 s) was refused by validation and the repair (5,431 in, 1,530 out, 20.5 s) was accepted | all valid, 15 to 16 s each, one call each |
| Behaviour check | 2 of 5 | 3 of 5 |
| Phase judge test | pass | pass |
Hint order (live_interviewer_poses_the_variant_and_serves_hints_in_order) |
fail: a second hint request answered without log_hint, on 3sum and on two-sum with the whiteboard |
pass |
| Uncertain speech | fail: wrong indices not challenged against the input | fail, the same way |
| Played candidates | fail, 5 lines | fail, 8 lines: hint requests answered without log_hint, and limits volunteered (3sum 100,000, coin-change 10,000, two-sum 10^9) |
The failures are model behaviour, not timeouts. Every report call stayed inside the 45 s limit, the longest at 24 s.
In a five-minute interview by hand on the same machine, with the interviewer also local, a report took one call and 17 s.
Worth trying
- Models: Qwen3-14B, Qwen3.5-9B, gpt-oss-20b, Mistral Small 3.x, Llama 3.x 8B, and larger quantizations of gemma-4-12b. Earlier trials on the interviewer side here: Qwen3.5-9B named the intended data structure unasked, and Qwen3-14B followed the hint rules best but did not fit fully in 16 GB beside the speech models.
- Hardware: 8 GB and 12 GB NVIDIA cards, AMD through ROCm or Vulkan, and Apple Silicon. On an M1 Pro, my estimate is that a report call takes 90 s or more, which would hit the 45 s limit; a measurement would settle it. If it does, the deadlines should probably be configurable for local setups, and the results here are what would size them.
Related: #105 (running CodeTrial on a local model), #110.
- 主要言語
- Rust
- スター
- 144
- フォーク
- 41
- 平均マージ
- 1日 7時間
- マージ済み PR(30日)
- 94
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
sysprog21/codetrial のほかの issue
-
Highlight the active line in the code editor対応中かも @ArthurArthurArthur0817 が 2 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
sysprog21/codetrial#247 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
Support configurable and randomized interviewer voices and accents対応中かも @MorganHo001 が 1 日前に担当しました。 オープン
sysprog21/codetrial#261 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
sysprog21/codetrial#258 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Add an example drawing area to coding interviews対応中かも @Chuyutseng が 1 日前に担当しました。 オープン
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
sysprog21/codetrial#255 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
Add indentation guides to the code editor対応中かも @sidney-Hung が 2 日前に担当しました。 オープン
sysprog21/codetrial#254 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
sysprog21/codetrial の issue をすべて見る
似ている issue
-
enhancement user-priority/P3
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 75/100
element-hq/lk-jwt-service#248 ·
メンテナーはふだん 1 日以内に返信
-
agent:triaged bug bughunt pm:pipenv priority:p1
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
SocketDev/socket-patch#1219 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
pact-foundation/pact-cli#154 ·
メンテナーはふだん 3 日以内に返信