Allow multi-thread streaming using the same weight loaded once in memory
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 55/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
調査の方向性
まず src/ggml_graph.cpp の28〜35行目にある mutex を確認し、次に parakeet-thread-local-backend.patch と concurrent_streams.py を調べてください。提供された Python コマンドを4つのスレッドで、提供されたモデルと音声ファイルを使って実行してください。1つのプロセス内で複数のストリームを重みを共有しながら同時に文字起こしでき、GPU システムでも動作すれば完了です。
索引モデルが issue の本文から書いたものです。
説明
I'm trying out this project for a live transcription usecase in a call. We use vosk at the moment (in https://github.com/nextcloud/live_transcription/) which supports loading the model weights once in the memory and then starting new threads with light-er recognize objects reading the same weights to process the output but keeping their own state and cache.
This is not possible at this moment due to a mutex https://github.com/mudler/parakeet.cpp/blob/e75de9b6b9b688fd293aa22f7e27aa724ea286f8/src/ggml_graph.cpp#L28-L35
One process can process only one stream at a time.
I'm not familiar with the code so did an experiment and asked AI if there is a possibility to work like llama.cpp here which uses slots to entertain parallel requests using the same loaded weights, and works with the same underlying ggml library.
It has successfully changed the code to make it possible for multiple threads to transcribe at the same time, in the same process by using a new Backend in each thread as opposed to one shared Backend + mutex guard.
The tests ran on an AMD CPU but theoritically should not cause issues with GPU systems.
Below are some reproduction steps and the patch:
parakeet-thread-local-backend.patch
concurrent_streams.py
python concurrent_streams.py --lib ./libparakeet.so --model ./nemotron-3.5-asr-streaming-0.6b-q8_0.gguf --threads 4 ./audio/en_9min_16k* --seconds 90
- 主要言語
- C++
- スター
- 831
- フォーク
- 101
- 平均マージ
- 1日 19時間
- マージ済み PR(30日)
- 26
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートなし
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mudler/parakeet.cpp のほかの issue
-
難易度 3/5 1〜2日 初心者へのやさしさ 58/100
mudler/parakeet.cpp#62 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 58/100
mudler/parakeet.cpp#61 ·
メンテナーはふだん 1 日以内に返信
-
Real streaming from a mic対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープン
難易度 3/5 1〜2日 初心者へのやさしさ 48/100
mudler/parakeet.cpp#60 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
mudler/parakeet.cpp#59 · コメント 3 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
mudler/parakeet.cpp#55 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
mudler/parakeet.cpp の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
NVIDIA/DeepStream#78 ·
-
Component: Python API
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
Vector35/binaryninja-api#8649 ·
メンテナーはふだん 3 日以内に返信
-
ai_p2 comp-parquet-reader-v3
難易度 2/5 半日 初心者へのやさしさ 66/100
ClickHouse/ClickHouse#124986 ·
メンテナーはふだん 1 日以内に返信
-
bug product: very_good_flutter_plugin
難易度 1/5 1〜3時間 初心者へのやさしさ 78/100
VeryGoodOpenSource/very_good_templates#654 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信