Dense FP32 (and FP16) matmul on the Android JNI tier: attention runs scalar today
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Área
- machine-learning, mobile, performance
Línea de trabajo
Start by reading fp32_matmul.c, fp16_matmul.c, and the JNI exports in native/skainet_jni.c; then trace how JniKernelProvider registers kernels and review JniKernelParityTest and KernelSupportMatrixTest. Add the requested JNI matmul variants and registrations, verify scalar parity including offset and strided operands, regenerate the support matrix, and measure the whisper-tiny encoder on a device. Done means bit-identical attention-shape results and at least 5× speedup on a Pixel 7 Pro/8a-class phone.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Context
A fully offline Android transcription app built on LiteRT (Whisper large-v3-turbo split encoder/decoder on OpenCL FP16, Parakeet TDT 0.6B, Silero VAD on ONNX Runtime) was evaluated as a SKaiNET consumer. Its author's own measurements on a Pixel 7 Pro are the bar: LiteRT CPU XNNPACK transcribes a 14 s memo in 82.2 s (RTF 5.9); the GPU path needs 34.4 s.
Since 0.50.0 the Android JNI tier (skainet-backend-jni-cpu) serves every GGML quant format from mapped weights on NEON, and 0.52.0 made the dispatch self-installing. That covers the weight side of a Whisper encoder (Q8_0 from whisper.cpp GGUFs, 874 MB). What it does not cover is the activation side.
Gap
The generated kernel support matrix (docs/.../reference/kernel-support-matrix.adoc, 0.54.0) lists Float32 and BFloat16 as scalar on Android, and native/skainet_jni.c exports only q40/q4k/q50/q51/q5k/q6k/q80 matmul entries plus the ternary gemv. A Whisper large-v3-turbo encoder issues, per layer, 20 heads of QKᵀ ([1500,64]×[64,1500]) and AV ([1500,1500]×[1500,64]) as FP32 activation × activation matmuls, 32 layers deep, plus two conv1d stem layers. On Android all of that runs through the scalar Kotlin path, so the quantized weight kernels cannot make the encoder fast on their own.
fp32_matmul.c already exists in skainet-backend-native-cpu (aarch64-verified, see #920) and the JVM reaches it through FFM; Android does not.
Scope
- JNI entries for dense FP32 matmul, including a batched variant and a transposed-right-operand variant (
A × Bᵀ) so attention does not pay a transpose copy per head, using the existingskainet_row_threadspool. - FP16 weight matmul on the JNI tier, since whisper.cpp ships F16 GGUFs and
fp16_matmul.cis in tree (#885 tracks the FFM side of the same kernel). -
JniKernelProviderregistrations soKernelDispatchselects them on Android at the native priority. - Parity tests against
ScalarFp32MatmulKernelinJniKernelParityTest, including offset/strided operands (the #1173 class of bug). -
KernelSupportMatrixTestregenerated:Float32andFloat16shownative-jnion Android. - Device measurement with the M2-A5 harness: whisper-tiny encoder before/after.
Acceptance
- Bit-identical results to the scalar path on device for the attention shapes above.
- whisper-tiny.en encoder on a Pixel 7 Pro / 8a class phone at least 5× faster than the scalar Android baseline.
Related
- #920 — NEON kernels to mobile (the JNI tier this extends)
- #885 — FP16 native kernel on the FFM tier
- #949 — the non-matmul Android overhead this issue does not address (see the companion issue on fused elementwise/norm kernels)
- Lenguaje dominante
- Kotlin
- Estrellas
- 52
- Forks
- 15
- Merge medio
- 1 d 15 h
- PR fusionados (30 d)
- 36
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de SKaiNET-developers/SKaiNET
-
coding good first issue size:xs skill:kotlin-core sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
SKaiNET-developers/SKaiNET#1323 ·
Los mantenedores suelen responder en 1 día
-
coding good first issue platform size:xs skill:js sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
SKaiNET-developers/SKaiNET#1232 ·
Los mantenedores suelen responder en 1 día
-
skeep tracking
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
SKaiNET-developers/SKaiNET#1331 ·
Los mantenedores suelen responder en 1 día
-
assessment size:s skill:review sub-issue
Dificultad 4/5 1-2 días Aptitud para principiantes 18/100
SKaiNET-developers/SKaiNET#1330 ·
Los mantenedores suelen responder en 1 día
-
documentation good first issue size:s skill:docs sub-issue
Dificultad 3/5 1-2 días Aptitud para principiantes 78/100
SKaiNET-developers/SKaiNET#1329 ·
Los mantenedores suelen responder en 1 día
Todos los issues de SKaiNET-developers/SKaiNET
Issues similares
-
[Submission] 抖音火山版Abiertosubmit-adaption submit-adaption-pre
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
BetterAndroid/android-notification-icon-project#744 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 64/100
utopia-rise/godot-jvm#1004 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
pedroSG94/RootEncoder#2213 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
[Feature]: AC charger voltageAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
dzid26/TeslaBatteryBLE#183 ·
Los mantenedores suelen responder en 1 día