Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Enable Nemotron-3-Diarization on MLX, CUDA, XNNPACK, and Vulkan

オープン
#23,131 コメント 0 件 リアクション 1 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
cpp, python, pytorch

調査の方向性

Start with the Nemotron-3-Diarization model card, checkpoint configuration, and Transformers integration linked in the issue, then compare them with the upstream NeMo reference and the NeMo-Speech.cpp option. Done requires export and native inference across MLX, CUDA, XNNPACK, and Vulkan, with documented commands, regression coverage, reference comparisons, and measured performance and delegation results.

索引モデルが issue の本文から書いたものです。

説明

enhancement module: cuda module: examples module: mlx module: vulkan module: xnnpack
🚀 The feature, motivation and pitch

Enable NVIDIA Nemotron-3-Diarization in ExecuTorch, with export and native inference examples for MLX, CUDA, XNNPACK, and Vulkan. Support offline recordings and streaming audio so applications can produce speaker timestamps locally for meetings, calls, and speech pipelines.

The model card describes a ~100M-parameter model supporting up to eight speakers, ordered by first arrival. It consumes 16 kHz mono audio, uses 128-bin mel features and a 31-layer Transformer with RoPE, stacks features to an 80 ms encoder frame rate, and uses a Conv1D upsampling head to produce per-speaker activity probabilities at a default 10 ms resolution. Streaming preserves speaker identities through an Arrival-Order Speaker Cache (AOSC) and FIFO context.

Proposed scope
  • Model loading and export: Add a reproducible example with pinned checkpoint and dependency revisions, a PyTorch reference, and backend selection. Export the neural computation to .pte plus any required backend artifacts. Define input/output shapes, supported dtypes, and bounded sequence lengths or padded chunk sizes. Keep the shared model representation portable across the four backends.
  • Native inference: Provide a C++ runner and audio-file CLI covering mel preprocessing, chunk scheduling, AOSC/FIFO updates, and conversion of activity probabilities into speaker/start/end segments. Expose streaming feed, reset, and final-flush behavior; document which work runs on the host. Inference should run without Python or NeMo installed.
  • Streaming and offline modes: Support the published 30.4 s offline-style configuration and 1.04 s, 0.64 s, and 0.32 s streaming presets. Preserve arrival-order speaker labels, look-ahead handling, padding masks, and timestamp alignment across chunks. These values are input-buffer latency, excluding compute time.
Backend workstreams
Backend Target Work to validate and enable
MLX Apple Silicon GPU Lower the encoder and speaker head through the MLX delegate; validate RoPE, masked attention, Conv1D, and varying chunk/cache lengths.
CUDA NVIDIA GPU Integrate the CUDA export/runtime path; validate attention, shape handling, supported precision, and host/device transfer costs. Document the tested GPU and CUDA requirements.
XNNPACK Arm and x86 CPU Establish a floating-point CPU baseline; inspect attention/matmul, normalization, and Conv1D delegation, and tune threading. Validate dynamic shapes or provide padded/static variants where needed.
Vulkan Supported Android and desktop GPUs Validate attention, RoPE, Conv1D, tensor layouts, shape changes, and device limits. Add lowering/kernel support as needed and document any partitions assigned to XNNPACK or portable kernels during export.

These are validation targets, not confirmed operator gaps. Record actual delegation coverage and remaining blockers for each backend. Establish floating-point correctness first; evaluate reduced precision and quantization separately against that baseline.

Acceptance criteria
  • Each backend has documented export/build/run commands and a successful end-to-end run on named hardware, with delegation coverage and fallback operators reported.
  • Compare preprocessing, per-frame probabilities, and final segments against the same pinned upstream reference. Report diarization error rate (DER) with the dataset, scoring settings, and agreed numerical/quality tolerances.
  • Cover silence, overlapping speech, speaker arrivals, up to eight speakers, short/final chunks, reset between recordings, cache rollover, and long recordings with bounded streaming memory.
  • Report model/artifact size, peak memory, real-time factor, and p50/p95 chunk processing latency at batch size 1, including preprocessing and state updates. Separate initialization/warm-up, input buffering, and steady-state compute; record hardware, precision, thread count, and streaming preset.
  • Add regression coverage for export, runtime correctness, and streaming state handling, with documentation of supported configurations and remaining limitations.
Alternatives

The upstream NeMo and Transformers implementations provide reference inference. NeMo-Speech.cpp provides another native deployment option. This request brings the model into ExecuTorch's runtime and delegate ecosystem.

Additional context

cc @SS-JIA @manuelcandales @digantdesai @cbilgin @GregoryComer @JakeStevens @iseeyuan @lucylq @helunwencser @tarun292 @kimishpatel @jackzhxng @Gasoonjia @metascroy

主要言語
Python
スター
5k
フォーク
1.2k
平均マージ
2日 12時間
マージ済み PR(30日)
588

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

pytorch/executorch のほかの issue

pytorch/executorch の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。