SOTA experiment: CacheBridge cross-model KV transfer for router handoffs
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 活発
- 技術スタック
- pytorch, rust
調査の方向性
ファイル、テスト、エントリポイントは指定されていません。まず、既存のモデルルーティングとアダプターベンチマークのエントリポイントを特定し、固定された条件下で、通常の re-prefill、以前の単純なマッピングのベースライン、CacheBridge 方式の head-matched mapping を比較します。必要なレイテンシ、品質、ストレージ、キャリブレーション、メモリ、帯域幅、コンテキスト長、方向、複数回の切り替えの各メトリクスが、ロールバックを維持したまま、指定された promotion gate を満たせば完了です。
索引モデルが issue の本文から書いたものです。
説明
Finding
CacheBridge: Efficient Cross-Model KV Cache Transfer (arXiv:2609.00891, submitted 2026-09-01) reports cross-model KV transfer intended to avoid full receiver prefill during model-family handoffs. The originating team reports 99.83% mean target retention on Qwen3 in one direction family, up to 3x faster mapper application, 8x lower mapper storage, matching with one tenth the calibration data, and a reported mapper construction reduction from 92.63s to 8.63s in a Qwen3 14B to 32B setting.
Evidence status: originating-team measured, GPU/model-family specific, not independently reproduced by RuV.
RuV opportunity
This maps to ruvLLM, RuVector, Cognitum model routing, MidStream, and distributed inference. It could make mid-session model escalation materially cheaper if cache transfer preserves quality.
Experiment
Do not replace prefill globally. Add an isolated adapter benchmark with:
- normal target re-prefill
- prior simple cross-model mapping baseline
- CacheBridge-style head-matched mapping
Pin model revisions, tokenizer, RoPE configuration, context lengths, CUDA, driver, PyTorch/serving runtime, GPU, calibration corpus, and random seeds.
Metrics
- target task accuracy / perplexity retention
- TTFT and p50/p95 handoff latency
- mapper construction time
- mapper storage
- calibration sample count
- GPU memory and transfer bandwidth
- failure by context length and transfer direction
- quality after multiple model switches
Falsification
Mandatory controls include just re-prefilling, prefix caching on the same model, and a simpler linear map. If handoff frequency is low or transferred quality falls materially, the added mapping layer should be rejected.
Promotion gate
At least 2x lower handoff prefill latency on a real RuV routed workload, target quality loss no greater than 1 absolute point, deterministic rollback to ordinary prefill, and no change to RVM authority or provenance semantics.
- 主要言語
- Rust
- スター
- 4.5k
- フォーク
- 603
- 平均マージ
- 1日 11時間
- マージ済み PR(30日)
- 56
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
ruvnet/RuVector のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
メンテナーはふだん 1 日以内に返信
ruvnet/RuVector の issue をすべて見る
似ている issue
-
type/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
stackabletech/kafka-operator#1033 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug good first issue needs testing
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 3 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
farion1231/cc-switch#7744 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
datafusion
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
apache/iceberg-rust#3297 ·
メンテナーはふだん 1 日以内に返信