Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

AutoModel passes the original source rate to already-resampled speaker segments

Aperta Adatta ai principianti
#3,762 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
72/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python, pytorch

Direzione di ricerca

Inizia in funasr/auto/auto_model.py, al confine del modello del parlante in AutoModel.generate, e traccia in che modo la frequenza del segmento ricampionato differisce dalla fs originale. Esegui lo script di riproduzione fornito per registrare il numero di campioni che entrano nell’estrazione delle caratteristiche del parlante. Il lavoro è completato quando il confine del parlante usa la frequenza effettiva del segmento, la riproduzione restituisce 24.000 campioni e i test di regressione mirati esistenti passano.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

🐛 Bug

On main commit 66d7a4c264a5993a2a63ed00c1f402c296ee521a, the VAD pipeline resamples source audio to the ASR frontend rate before slicing and speaker chunking. The ASR boundary explicitly uses the segment rate, but the speaker boundary forwards the original generate(..., fs=...) value. CAMPPlus uses this as its source rate and resamples the waveform again.

A 200 ms synthetic 48 kHz input becomes a 16 kHz segment, padded by sv_chunk to 1.5 seconds (24,000 samples). Speaker feature extraction instead receives 8,000 samples. With an 8 kHz source override the same padded segment becomes 48,000 samples. An incompatible spk_kwargs.fs can also override the actual segment rate.

PR #3751 fixed the VAD→ASR rate propagation and retained the original rate across input batches; this report concerns the remaining speaker boundary.

To Reproduce

Use a source checkout at the SHA above and a local Python 3.12 environment. The following script uses actual AutoModel generation, audio resampling, speaker chunking and CAMPPlus inference, with small inference stand-ins and a feature-capture boundary. It downloads no model weights.

Tested install:

python -m pip install torch==2.10.0 torchaudio==2.10.0 numpy==1.26.4 transformers==4.51.3 einops pytorch-wpe -e .

Save the code below as reproduce_speaker_rate.py, then run:

HF_HUB_OFFLINE=1 MODELSCOPE_OFFLINE=1 OMP_NUM_THREADS=1 python reproduce_speaker_rate.py

Code sample

"""Weights-free public-entry reproduction; real loading/resampling and CAMPPlus entry."""

from types import SimpleNamespace
from unittest.mock import patch

import numpy as np
import torch

from funasr.auto.auto_model import AutoModel
from funasr.models.campplus import model as campplus_model


class Probe(torch.nn.Module):
    def __init__(self, role):
        super().__init__()
        self.anchor = torch.nn.Parameter(torch.zeros(1))
        self.role = role

    def inference(self, data_in, key, **kwargs):
        if self.role == 'speaker':
            print('speaker received fs:', kwargs.get('fs', 16000))
            return campplus_model.CAMPPlus.inference(self, data_in, key=key, **kwargs)
        results = [
            {'key': k, 'value': [[0, 200]]} if self.role == 'vad'
            else {'key': k, 'text': 'probe'} for k in key
        ]
        return results, {'batch_data_time': 0.2}

    def forward(self, features):
        return torch.ones(features.shape[0], 4)


captured = []


def capture_features(audio):
    captured.extend(len(sample) for sample in audio)
    return torch.zeros(len(audio), 20, 80), [20] * len(audio), captured[-len(audio):]


wrapper = AutoModel.__new__(AutoModel)
wrapper.model = Probe('asr')
wrapper.vad_model = Probe('vad')
wrapper.spk_model = Probe('speaker')
wrapper.punc_model = None
wrapper.kwargs = {
    'device': 'cpu', 'disable_pbar': True, 'return_spk_res': False,
    'frontend': SimpleNamespace(fs=16000), 'ncpu': 1,
}
wrapper.vad_kwargs = {'device': 'cpu'}
wrapper.spk_kwargs = {'device': 'cpu'}
wrapper._store_base_configs()
source_rate = 48000
audio = np.sin(2 * np.pi * 440 * np.arange(source_rate // 5) / source_rate).astype(np.float32)
with patch.object(campplus_model, 'extract_feature', capture_features):
    wrapper.generate(input=audio, fs=source_rate)
print('sample counts entering speaker feature extraction:', captured)
print('expected: [24000] (1.5 seconds padded at 16 kHz)')
assert captured == [24000], captured

Expected behavior

The speaker model receives the actual rate of its input segments (fs=16000 in this pipeline) and feature extraction sees 24,000 samples. An explicit source rate still reaches VAD/source loading and constructor configuration remains available for the next call.

Error logs

speaker received fs: 48000
sample counts entering speaker feature extraction: [8000]
expected: [24000] (1.5 seconds padded at 16 kHz)
AssertionError: [8000]

The proposed one-line boundary fix produces fs: 16000 and [24000] on the same script. New focused regressions fail 21 cases on the baseline while 8 controls pass; the fixed source passes all 29. Tests cover arrays/explicit PCM, single/batch inputs, constructor/runtime precedence, speaker config conflicts and subsequent calls. No recognition or diarization quality claim is made from these stand-in models.

Environment

  • OS: macOS 26.6.2 arm64
  • Python: 3.12.12
  • FunASR: source 1.4.16, SHA above
  • ModelScope: 1.40.1
  • PyTorch / torchaudio: 2.10.0 / 2.10.0
  • NumPy: 1.26.4; transformers: 4.51.3; tokenizers: 0.21.4
  • Install method: editable source in a task-local virtual environment
  • Device: CPU; no GPU/CUDA/Docker execution

Audio details

Synthetic mono float32 440 Hz sine, 200 ms, source rate 48 kHz; additional regressions use 8/16 kHz and explicit signed PCM16. Speaker chunking pads it to 1.5 seconds. There is no speech, real speaker identity or background audio.

Prepared with AI assistance; this is a current-source reproduction, not a recovered historical patch.

Lingua principale
Python
Stelle
20.6k
Fork
2.1k
Merge medio
18h 3m
PR unite (30g)
88

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di modelscope/FunASR

Tutte le issue di modelscope/FunASR

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.