Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

SOTA experiment: CacheBridge cross-model KV transfer for router handoffs

オープン
#956 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
活発
技術スタック
pytorch, rust

調査の方向性

ファイル、テスト、エントリポイントは指定されていません。まず、既存のモデルルーティングとアダプターベンチマークのエントリポイントを特定し、固定された条件下で、通常の re-prefill、以前の単純なマッピングのベースライン、CacheBridge 方式の head-matched mapping を比較します。必要なレイテンシ、品質、ストレージ、キャリブレーション、メモリ、帯域幅、コンテキスト長、方向、複数回の切り替えの各メトリクスが、ロールバックを維持したまま、指定された promotion gate を満たせば完了です。

索引モデルが issue の本文から書いたものです。

説明

Finding

CacheBridge: Efficient Cross-Model KV Cache Transfer (arXiv:2609.00891, submitted 2026-09-01) reports cross-model KV transfer intended to avoid full receiver prefill during model-family handoffs. The originating team reports 99.83% mean target retention on Qwen3 in one direction family, up to 3x faster mapper application, 8x lower mapper storage, matching with one tenth the calibration data, and a reported mapper construction reduction from 92.63s to 8.63s in a Qwen3 14B to 32B setting.

Evidence status: originating-team measured, GPU/model-family specific, not independently reproduced by RuV.

RuV opportunity

This maps to ruvLLM, RuVector, Cognitum model routing, MidStream, and distributed inference. It could make mid-session model escalation materially cheaper if cache transfer preserves quality.

Experiment

Do not replace prefill globally. Add an isolated adapter benchmark with:

  1. normal target re-prefill
  2. prior simple cross-model mapping baseline
  3. CacheBridge-style head-matched mapping

Pin model revisions, tokenizer, RoPE configuration, context lengths, CUDA, driver, PyTorch/serving runtime, GPU, calibration corpus, and random seeds.

Metrics

  • target task accuracy / perplexity retention
  • TTFT and p50/p95 handoff latency
  • mapper construction time
  • mapper storage
  • calibration sample count
  • GPU memory and transfer bandwidth
  • failure by context length and transfer direction
  • quality after multiple model switches

Falsification

Mandatory controls include just re-prefilling, prefix caching on the same model, and a simpler linear map. If handoff frequency is low or transferred quality falls materially, the added mapping layer should be rejected.

Promotion gate

At least 2x lower handoff prefill latency on a real RuV routed workload, target quality loss no greater than 1 absolute point, deterministic rollback to ordinary prefill, and no change to RVM authority or provenance semantics.

主要言語
Rust
スター
4.5k
フォーク
603
平均マージ
1日 11時間
マージ済み PR(30日)
56

環境構築

このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

ruvnet/RuVector のほかの issue

ruvnet/RuVector の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。