Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Allow multi-thread streaming using the same weight loaded once in memory

Aperta
#68 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
55/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp, python

Direzione di ricerca

Inizia esaminando il mutex in src/ggml_graph.cpp alle righe 28-35, quindi ispeziona parakeet-thread-local-backend.patch e concurrent_streams.py. Esegui il comando Python fornito con quattro thread e con il modello e i file audio forniti. Il lavoro è completato quando più stream possono essere trascritti contemporaneamente in un unico processo condividendo i pesi, anche su sistemi GPU.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

I'm trying out this project for a live transcription usecase in a call. We use vosk at the moment (in https://github.com/nextcloud/live_transcription/) which supports loading the model weights once in the memory and then starting new threads with light-er recognize objects reading the same weights to process the output but keeping their own state and cache.
This is not possible at this moment due to a mutex https://github.com/mudler/parakeet.cpp/blob/e75de9b6b9b688fd293aa22f7e27aa724ea286f8/src/ggml_graph.cpp#L28-L35
One process can process only one stream at a time.

I'm not familiar with the code so did an experiment and asked AI if there is a possibility to work like llama.cpp here which uses slots to entertain parallel requests using the same loaded weights, and works with the same underlying ggml library.
It has successfully changed the code to make it possible for multiple threads to transcribe at the same time, in the same process by using a new Backend in each thread as opposed to one shared Backend + mutex guard.
The tests ran on an AMD CPU but theoritically should not cause issues with GPU systems.

Below are some reproduction steps and the patch:

parakeet-thread-local-backend.patch
concurrent_streams.py

python concurrent_streams.py --lib ./libparakeet.so --model ./nemotron-3.5-asr-streaming-0.6b-q8_0.gguf --threads 4 ./audio/en_9min_16k* --seconds 90
Lingua principale
C++
Stelle
786
Fork
93
Merge medio
9g 19h
PR unite (30g)
4

Preparare l'ambiente

Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/parakeet.cpp

Tutte le issue di mudler/parakeet.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.