[enhancement] Variance-normalized KV-cache quantization (KVarN) & KV cache precision tail
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- cpp
- Domain
- machine-learning
Research direction
The issue names Strata's KV cache and KVarN types and tail in beellama.cpp, but gives no files, tests, or implementation requirements. Start by locating Strata's KV-cache quantization entry points and comparing them with the referenced KVarN implementation; first clarify the desired behavior and feasibility, then define tests that would demonstrate support.
Written by the indexing model from the issue text.
Description
I'm comming from beellama.cpp and use the KVarN kv-cache types & tail.
Maybe you have time to take a look at this, if it's possible to implement in Strata, if you consider it as valueadding?
Great work sofar, many thanks!
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Possibly taken @nekomario28 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Possibly taken @nekomario28 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Niko1221/Strata#1463 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Niko1221/Strata#1322 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
Status: Awaiting triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
espressif/arduino-esp32#12984 ·
Maintainers usually reply within 1 day
-
torch_ops/logprob.cu does not compile with the serving container's nvcc (13.3.73); check_torch_ops.py cannot run as shippedPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 Under an hour Newbie friendliness 72/100
ashhart/TensorFold#535 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NVIDIA/cuCollections#858 ·
Maintainers usually reply within 2 days