Dense FP32 (and FP16) matmul on the Android JNI tier: attention runs scalar today
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Active
- Domain
- machine-learning, mobile, performance
Research direction
Start by reading fp32_matmul.c, fp16_matmul.c, and the JNI exports in native/skainet_jni.c; then trace how JniKernelProvider registers kernels and review JniKernelParityTest and KernelSupportMatrixTest. Add the requested JNI matmul variants and registrations, verify scalar parity including offset and strided operands, regenerate the support matrix, and measure the whisper-tiny encoder on a device. Done means bit-identical attention-shape results and at least 5× speedup on a Pixel 7 Pro/8a-class phone.
Written by the indexing model from the issue text.
Description
Context
A fully offline Android transcription app built on LiteRT (Whisper large-v3-turbo split encoder/decoder on OpenCL FP16, Parakeet TDT 0.6B, Silero VAD on ONNX Runtime) was evaluated as a SKaiNET consumer. Its author's own measurements on a Pixel 7 Pro are the bar: LiteRT CPU XNNPACK transcribes a 14 s memo in 82.2 s (RTF 5.9); the GPU path needs 34.4 s.
Since 0.50.0 the Android JNI tier (skainet-backend-jni-cpu) serves every GGML quant format from mapped weights on NEON, and 0.52.0 made the dispatch self-installing. That covers the weight side of a Whisper encoder (Q8_0 from whisper.cpp GGUFs, 874 MB). What it does not cover is the activation side.
Gap
The generated kernel support matrix (docs/.../reference/kernel-support-matrix.adoc, 0.54.0) lists Float32 and BFloat16 as scalar on Android, and native/skainet_jni.c exports only q40/q4k/q50/q51/q5k/q6k/q80 matmul entries plus the ternary gemv. A Whisper large-v3-turbo encoder issues, per layer, 20 heads of QKᵀ ([1500,64]×[64,1500]) and AV ([1500,1500]×[1500,64]) as FP32 activation × activation matmuls, 32 layers deep, plus two conv1d stem layers. On Android all of that runs through the scalar Kotlin path, so the quantized weight kernels cannot make the encoder fast on their own.
fp32_matmul.c already exists in skainet-backend-native-cpu (aarch64-verified, see #920) and the JVM reaches it through FFM; Android does not.
Scope
- JNI entries for dense FP32 matmul, including a batched variant and a transposed-right-operand variant (
A × Bᵀ) so attention does not pay a transpose copy per head, using the existingskainet_row_threadspool. - FP16 weight matmul on the JNI tier, since whisper.cpp ships F16 GGUFs and
fp16_matmul.cis in tree (#885 tracks the FFM side of the same kernel). -
JniKernelProviderregistrations soKernelDispatchselects them on Android at the native priority. - Parity tests against
ScalarFp32MatmulKernelinJniKernelParityTest, including offset/strided operands (the #1173 class of bug). -
KernelSupportMatrixTestregenerated:Float32andFloat16shownative-jnion Android. - Device measurement with the M2-A5 harness: whisper-tiny encoder before/after.
Acceptance
- Bit-identical results to the scalar path on device for the attention shapes above.
- whisper-tiny.en encoder on a Pixel 7 Pro / 8a class phone at least 5× faster than the scalar Android baseline.
Related
- #920 — NEON kernels to mobile (the JNI tier this extends)
- #885 — FP16 native kernel on the FFM tier
- #949 — the non-matmul Android overhead this issue does not address (see the companion issue on fused elementwise/norm kernels)
- Dominant language
- Kotlin
- Stars
- 52
- Forks
- 15
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 36
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from SKaiNET-developers/SKaiNET
-
coding good first issue size:xs skill:kotlin-core sub-issue
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
SKaiNET-developers/SKaiNET#1323 ·
Maintainers usually reply within 1 day
-
coding good first issue platform size:xs skill:js sub-issue
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
SKaiNET-developers/SKaiNET#1232 ·
Maintainers usually reply within 1 day
-
skeep tracking
Difficulty 5/5 Over a week Newbie friendliness 15/100
SKaiNET-developers/SKaiNET#1331 ·
Maintainers usually reply within 1 day
-
assessment size:s skill:review sub-issue
Difficulty 4/5 1-2 days Newbie friendliness 18/100
SKaiNET-developers/SKaiNET#1330 ·
Maintainers usually reply within 1 day
-
documentation good first issue size:s skill:docs sub-issue
Difficulty 3/5 1-2 days Newbie friendliness 78/100
SKaiNET-developers/SKaiNET#1329 ·
Maintainers usually reply within 1 day
All issues in SKaiNET-developers/SKaiNET
Similar issues
-
[Submission] 抖音火山版Opensubmit-adaption submit-adaption-pre
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
BetterAndroid/android-notification-icon-project#744 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
utopia-rise/godot-jvm#1004 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
pedroSG94/RootEncoder#2213 ·
Maintainers usually reply within 2 days
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
dzid26/TeslaBatteryBLE#183 ·
Maintainers usually reply within 1 day