Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

AutoModel passes the original source rate to already-resampled speaker segments

未关闭 适合新手
#3,762 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

@hulkbig 已经在做这个了。

开始于 2026年10月5日。

  • #3763 来自 @hulkbig —— 未关闭

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
72/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python, pytorch

调研方向

从 funasr/auto/auto_model.py 中 AutoModel.generate 的说话人模型边界开始,追踪重采样片段的采样率与原始 fs 有何不同。运行提供的复现脚本,记录进入说话人特征提取的样本数。完成标准是说话人边界使用片段的实际采样率,复现结果为24,000个样本,并且现有的定向回归测试通过。

由索引模型根据 Issue 内容生成。

描述

🐛 Bug

On main commit 66d7a4c264a5993a2a63ed00c1f402c296ee521a, the VAD pipeline resamples source audio to the ASR frontend rate before slicing and speaker chunking. The ASR boundary explicitly uses the segment rate, but the speaker boundary forwards the original generate(..., fs=...) value. CAMPPlus uses this as its source rate and resamples the waveform again.

A 200 ms synthetic 48 kHz input becomes a 16 kHz segment, padded by sv_chunk to 1.5 seconds (24,000 samples). Speaker feature extraction instead receives 8,000 samples. With an 8 kHz source override the same padded segment becomes 48,000 samples. An incompatible spk_kwargs.fs can also override the actual segment rate.

PR #3751 fixed the VAD→ASR rate propagation and retained the original rate across input batches; this report concerns the remaining speaker boundary.

To Reproduce

Use a source checkout at the SHA above and a local Python 3.12 environment. The following script uses actual AutoModel generation, audio resampling, speaker chunking and CAMPPlus inference, with small inference stand-ins and a feature-capture boundary. It downloads no model weights.

Tested install:

python -m pip install torch==2.10.0 torchaudio==2.10.0 numpy==1.26.4 transformers==4.51.3 einops pytorch-wpe -e .

Save the code below as reproduce_speaker_rate.py, then run:

HF_HUB_OFFLINE=1 MODELSCOPE_OFFLINE=1 OMP_NUM_THREADS=1 python reproduce_speaker_rate.py

Code sample

"""Weights-free public-entry reproduction; real loading/resampling and CAMPPlus entry."""

from types import SimpleNamespace
from unittest.mock import patch

import numpy as np
import torch

from funasr.auto.auto_model import AutoModel
from funasr.models.campplus import model as campplus_model


class Probe(torch.nn.Module):
    def __init__(self, role):
        super().__init__()
        self.anchor = torch.nn.Parameter(torch.zeros(1))
        self.role = role

    def inference(self, data_in, key, **kwargs):
        if self.role == 'speaker':
            print('speaker received fs:', kwargs.get('fs', 16000))
            return campplus_model.CAMPPlus.inference(self, data_in, key=key, **kwargs)
        results = [
            {'key': k, 'value': [[0, 200]]} if self.role == 'vad'
            else {'key': k, 'text': 'probe'} for k in key
        ]
        return results, {'batch_data_time': 0.2}

    def forward(self, features):
        return torch.ones(features.shape[0], 4)


captured = []


def capture_features(audio):
    captured.extend(len(sample) for sample in audio)
    return torch.zeros(len(audio), 20, 80), [20] * len(audio), captured[-len(audio):]


wrapper = AutoModel.__new__(AutoModel)
wrapper.model = Probe('asr')
wrapper.vad_model = Probe('vad')
wrapper.spk_model = Probe('speaker')
wrapper.punc_model = None
wrapper.kwargs = {
    'device': 'cpu', 'disable_pbar': True, 'return_spk_res': False,
    'frontend': SimpleNamespace(fs=16000), 'ncpu': 1,
}
wrapper.vad_kwargs = {'device': 'cpu'}
wrapper.spk_kwargs = {'device': 'cpu'}
wrapper._store_base_configs()
source_rate = 48000
audio = np.sin(2 * np.pi * 440 * np.arange(source_rate // 5) / source_rate).astype(np.float32)
with patch.object(campplus_model, 'extract_feature', capture_features):
    wrapper.generate(input=audio, fs=source_rate)
print('sample counts entering speaker feature extraction:', captured)
print('expected: [24000] (1.5 seconds padded at 16 kHz)')
assert captured == [24000], captured

Expected behavior

The speaker model receives the actual rate of its input segments (fs=16000 in this pipeline) and feature extraction sees 24,000 samples. An explicit source rate still reaches VAD/source loading and constructor configuration remains available for the next call.

Error logs

speaker received fs: 48000
sample counts entering speaker feature extraction: [8000]
expected: [24000] (1.5 seconds padded at 16 kHz)
AssertionError: [8000]

The proposed one-line boundary fix produces fs: 16000 and [24000] on the same script. New focused regressions fail 21 cases on the baseline while 8 controls pass; the fixed source passes all 29. Tests cover arrays/explicit PCM, single/batch inputs, constructor/runtime precedence, speaker config conflicts and subsequent calls. No recognition or diarization quality claim is made from these stand-in models.

Environment

  • OS: macOS 26.6.2 arm64
  • Python: 3.12.12
  • FunASR: source 1.4.16, SHA above
  • ModelScope: 1.40.1
  • PyTorch / torchaudio: 2.10.0 / 2.10.0
  • NumPy: 1.26.4; transformers: 4.51.3; tokenizers: 0.21.4
  • Install method: editable source in a task-local virtual environment
  • Device: CPU; no GPU/CUDA/Docker execution

Audio details

Synthetic mono float32 440 Hz sine, 200 ms, source rate 48 kHz; additional regressions use 8/16 kHz and explicit signed PCM16. Speaker chunking pads it to 1.5 seconds. There is no speech, real speaker identity or background audio.

Prepared with AI assistance; this is a current-source reproduction, not a recovered historical patch.

主要语言
Python
星标
20.6k
派生
2.1k
平均合并
18 小时 3 分钟
30 天内合并 PR
88

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

modelscope/FunASR 的其他 Issue

查看 modelscope/FunASR 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。