Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

fix(cache): bound the turbo4 sliding append to visible_len

Aperta
#1,793 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
52/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
rust

Direzione di ricerca

Inizia in src/lib/mlxcel-core/src/cache.rs confrontando update_turbo4_concat con update_concat, visible_fp16_prefix_for_concat e trim. Esegui i test della cache denominati di Gemma 4 e Gemma 3, aggiungendo i casi turbo e rollback descritti nei criteri di accettazione; il lavoro è completato quando le chiavi restituite corrispondono alla lunghezza attiva senza avvisi sulla mask. Verifica il percorso di tracciamento dei logits e concludi con i controlli di test del workspace, clippy e fmt.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

area:core area:inference priority:medium status:ready type:bug

Problem

With --kv-cache-mode fp16+turbo4 (KVCacheMode::Turbo4Asym), the multi-token append of Gemma 4's sliding caches (update_turbo4_concat) concatenates the whole physical key buffer, not the visible_len() prior keys the FP16 update_concat takes. PR #1752 and PR #1776 sized sliding masks from visible_len() to match FP16; the turbo append was not changed. After a decode step grows the buffer, or a speculative trim rewinds offset and idx without moving data, it returns physical + l keys against an offset + l mask and attends zero-filled growth slots and rolled-back draft keys.

Nothing aborts: trim_mask_to_keys discards the too-short mask with a warning (one per sliding layer per append) and causal_attention rebuilds the band. Whether the extra keys change the attention output is unmeasured. It is reachable with Gemma 4 MTP on a turbo cache: enable_speculative_buffer refuses non-FP16 caches, Gemma 4 only logs that, and the next verify block after a rollback hits this. Found during review of PR #1776.

Evidence

  • src/lib/mlxcel-core/src/cache.rs: update_turbo4_concat :5003 (physical length :5028-5031, whole-buffer concat :5033-5039); FP16 update_concat :4686 via visible_fp16_prefix_for_concat :4752; turbo growth :5187-5202; trim :5454-5466; FP16-only refusal in enable_speculative_buffer :4425-4431.
  • src/models/gemma4.rs: per-layer mode in make_caches_with_modes :3801-3820; refusal only warned at :5176 and src/models/gemma4_mtp_target.rs:1717; trim_mask_to_keys discard branch :2560-2589.
  • The first_cache_live_len doc comment PR #1776 added (gemma4.rs:3923-3946) defines the live length as the prior keys the next append keeps, true only for FP16, and calls the turbo mask "cropped". It is shorter than the keys, so it is discarded.

Proposed fix

Take the visible_len() prior keys exactly as FP16 does, applying the same slices to v_packed, v_norms and v_rescale (generalize visible_fp16_prefix_for_concat rather than copying it). "Exactly" includes chronological order after a wrap (logical_start) and keeping max_size - 1 prior keys plus all new ones (#678), where turbo now clamps the return to max_size. Correct the doc comment.

Acceptance criteria

  • Turbo variants of first_cache_live_len_sliding_matches_returned_keys (gemma4.rs:6095) and live_len_matches_the_keys_a_multi_token_append_returns (src/models/gemma3_tests.rs:698), plus a trim rollback case, assert that the live length equals the returned key axis, and fail with the fix reverted.
  • No mask/key mismatch warning on a Gemma 4 MTP run with fp16+turbo4.
  • A teacher-forced logit trace with and without the fix on gemma-4-26b-a4b-it-4bit records whether decided positions change, per docs/benchmarks.md "Judging a change that moves the numbers".

Verification

examples/logit_trace cannot reach this as it stands: it builds FP16 caches via model.make_caches() (examples/logit_trace.rs:241) and runs only multi-token forwards (:253, :257). The trace arm needs fp16+turbo4 caches and a single-token step before the traced append, and the reverted arm must log the warning to prove the input reached the branch. Then the workspace gate (cargo test --workspace --profile test-fast --features metal,accelerate, clippy, fmt).

Lingua principale
Rust
Stelle
471
Fork
55
Merge medio
9h 36m
PR unite (30g)
289

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di lablup/mlxcel

Tutte le issue di lablup/mlxcel

Issue simili

Altre issue su Rust

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.