Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Native Silero VAD model on Lstm with weights from the upstream ONNX file

Abierto
#1,278 1 comentario 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
30/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
android, kotlin

Línea de trabajo

Start by reading the existing Lstm state-step API from #823 and the ONNX weight reader in skainet-io-onnx; the issue does not name specific files. Implement the Silero graph with ctx.ops.*, load weights from the pinned upstream ONNX initializers, and add the streaming adapter. Done means parity within 1e-4 over at least 1,000 frames, Android performance under 1 ms per frame, and a Python-side ground-truth entry.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement model skill:numerics

Context

Silero VAD (v6.x, MIT) is the de-facto on-device voice activity detector. A SKaiNET consumer currently has to ship ONNX Runtime only for it, next to whatever runs the ASR model, which costs a second native runtime, an extra ~2.3 MB model on a separate loader, and a minSdk floor set by ORT rather than by SKaiNET.

The model is small: an STFT front end baked into the graph as conv weights, a conv encoder, one LSTM cell with a carried [2,1,128] state, and a dense sigmoid head. Input per call is a 64-sample context prepended to a 512-sample frame at 16 kHz. Lstm with LstmState(h, c) and an explicit step already exists in skainet-lang-core (#823), which was added with exactly this kind of streaming use in mind.

Scope

  • SileroVad module (skainet-models or a new skainet-model-vad): the graph expressed with ctx.ops.* only, so it runs eagerly and lowers through the tape path like every other SKaiNET model. Reuse Lstm.
  • Weights loaded from the initializers of the pinned upstream silero_vad.onnx through skainet-io-onnx's weight reader (no new file format; upstream stays the source of truth). Load the STFT basis as weights rather than re-deriving it so the model matches bit for bit.
  • Streaming API: process(frame: FloatArray): Float returning speech probability, reset(), and the state carried across calls; a VoiceActivityDetector-shaped adapter for the audio-side libraries.
  • Parity test: probabilities within 1e-4 of ONNX Runtime over ≥ 1000 consecutive frames from three clips including state carry-over, checked against a committed golden (ORT is a test-time dependency only).
  • Android measurement: < 1 ms per frame on one thread with the JNI tier.
  • Ground-truth entry so the Python-side suite (#985) covers the op set.

Non-goals

  • Importing the Silero graph through the ONNX graph importer. That needs LSTM/Gather/Unsqueeze/Where/If converters and is tracked separately as the verification path; the hand-written module is the shipping path.

Related

  • #823 — Lstm layer with explicit state step API
  • #230 — ONNX I/O module MVP
Lenguaje dominante
Kotlin
Estrellas
52
Forks
15
Merge medio
1 d 15 h
PR fusionados (30 d)
36

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de SKaiNET-developers/SKaiNET

Todos los issues de SKaiNET-developers/SKaiNET

Issues similares

Más issues de Kotlin

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.