Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Dense FP32 (and FP16) matmul on the Android JNI tier: attention runs scalar today

Open
#1,275 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
android, c, kotlin

Research direction

Start by reading fp32_matmul.c, fp16_matmul.c, and the JNI exports in native/skainet_jni.c; then trace how JniKernelProvider registers kernels and review JniKernelParityTest and KernelSupportMatrixTest. Add the requested JNI matmul variants and registrations, verify scalar parity including offset and strided operands, regenerate the support matrix, and measure the whisper-tiny encoder on a device. Done means bit-identical attention-shape results and at least 5× speedup on a Pixel 7 Pro/8a-class phone.

Written by the indexing model from the issue text.

Description

compute-backend enhancement skill:android skill:native

Context

A fully offline Android transcription app built on LiteRT (Whisper large-v3-turbo split encoder/decoder on OpenCL FP16, Parakeet TDT 0.6B, Silero VAD on ONNX Runtime) was evaluated as a SKaiNET consumer. Its author's own measurements on a Pixel 7 Pro are the bar: LiteRT CPU XNNPACK transcribes a 14 s memo in 82.2 s (RTF 5.9); the GPU path needs 34.4 s.

Since 0.50.0 the Android JNI tier (skainet-backend-jni-cpu) serves every GGML quant format from mapped weights on NEON, and 0.52.0 made the dispatch self-installing. That covers the weight side of a Whisper encoder (Q8_0 from whisper.cpp GGUFs, 874 MB). What it does not cover is the activation side.

Gap

The generated kernel support matrix (docs/.../reference/kernel-support-matrix.adoc, 0.54.0) lists Float32 and BFloat16 as scalar on Android, and native/skainet_jni.c exports only q40/q4k/q50/q51/q5k/q6k/q80 matmul entries plus the ternary gemv. A Whisper large-v3-turbo encoder issues, per layer, 20 heads of QKᵀ ([1500,64]×[64,1500]) and AV ([1500,1500]×[1500,64]) as FP32 activation × activation matmuls, 32 layers deep, plus two conv1d stem layers. On Android all of that runs through the scalar Kotlin path, so the quantized weight kernels cannot make the encoder fast on their own.

fp32_matmul.c already exists in skainet-backend-native-cpu (aarch64-verified, see #920) and the JVM reaches it through FFM; Android does not.

Scope

  • JNI entries for dense FP32 matmul, including a batched variant and a transposed-right-operand variant (A × Bᵀ) so attention does not pay a transpose copy per head, using the existing skainet_row_threads pool.
  • FP16 weight matmul on the JNI tier, since whisper.cpp ships F16 GGUFs and fp16_matmul.c is in tree (#885 tracks the FFM side of the same kernel).
  • JniKernelProvider registrations so KernelDispatch selects them on Android at the native priority.
  • Parity tests against ScalarFp32MatmulKernel in JniKernelParityTest, including offset/strided operands (the #1173 class of bug).
  • KernelSupportMatrixTest regenerated: Float32 and Float16 show native-jni on Android.
  • Device measurement with the M2-A5 harness: whisper-tiny encoder before/after.

Acceptance

  • Bit-identical results to the scalar path on device for the attention shapes above.
  • whisper-tiny.en encoder on a Pixel 7 Pro / 8a class phone at least 5× faster than the scalar Android baseline.

Related

  • #920 — NEON kernels to mobile (the JNI tier this extends)
  • #885 — FP16 native kernel on the FFM tier
  • #949 — the non-matmul Android overhead this issue does not address (see the companion issue on fused elementwise/norm kernels)
Dominant language
Kotlin
Stars
52
Forks
15
Avg merge
1d 15h
Merged PRs (30d)
36

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SKaiNET-developers/SKaiNET

All issues in SKaiNET-developers/SKaiNET

Similar issues

More Kotlin issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.