AutoModel passes the original source rate to already-resampled speaker segments
维护者通常 1 天内回复
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 72/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
调研方向
从 funasr/auto/auto_model.py 中 AutoModel.generate 的说话人模型边界开始,追踪重采样片段的采样率与原始 fs 有何不同。运行提供的复现脚本,记录进入说话人特征提取的样本数。完成标准是说话人边界使用片段的实际采样率,复现结果为24,000个样本,并且现有的定向回归测试通过。
由索引模型根据 Issue 内容生成。
描述
🐛 Bug
On main commit 66d7a4c264a5993a2a63ed00c1f402c296ee521a, the VAD pipeline resamples source audio to the ASR frontend rate before slicing and speaker chunking. The ASR boundary explicitly uses the segment rate, but the speaker boundary forwards the original generate(..., fs=...) value. CAMPPlus uses this as its source rate and resamples the waveform again.
A 200 ms synthetic 48 kHz input becomes a 16 kHz segment, padded by sv_chunk to 1.5 seconds (24,000 samples). Speaker feature extraction instead receives 8,000 samples. With an 8 kHz source override the same padded segment becomes 48,000 samples. An incompatible spk_kwargs.fs can also override the actual segment rate.
PR #3751 fixed the VAD→ASR rate propagation and retained the original rate across input batches; this report concerns the remaining speaker boundary.
To Reproduce
Use a source checkout at the SHA above and a local Python 3.12 environment. The following script uses actual AutoModel generation, audio resampling, speaker chunking and CAMPPlus inference, with small inference stand-ins and a feature-capture boundary. It downloads no model weights.
Tested install:
python -m pip install torch==2.10.0 torchaudio==2.10.0 numpy==1.26.4 transformers==4.51.3 einops pytorch-wpe -e .
Save the code below as reproduce_speaker_rate.py, then run:
HF_HUB_OFFLINE=1 MODELSCOPE_OFFLINE=1 OMP_NUM_THREADS=1 python reproduce_speaker_rate.py
Code sample
"""Weights-free public-entry reproduction; real loading/resampling and CAMPPlus entry."""
from types import SimpleNamespace
from unittest.mock import patch
import numpy as np
import torch
from funasr.auto.auto_model import AutoModel
from funasr.models.campplus import model as campplus_model
class Probe(torch.nn.Module):
def __init__(self, role):
super().__init__()
self.anchor = torch.nn.Parameter(torch.zeros(1))
self.role = role
def inference(self, data_in, key, **kwargs):
if self.role == 'speaker':
print('speaker received fs:', kwargs.get('fs', 16000))
return campplus_model.CAMPPlus.inference(self, data_in, key=key, **kwargs)
results = [
{'key': k, 'value': [[0, 200]]} if self.role == 'vad'
else {'key': k, 'text': 'probe'} for k in key
]
return results, {'batch_data_time': 0.2}
def forward(self, features):
return torch.ones(features.shape[0], 4)
captured = []
def capture_features(audio):
captured.extend(len(sample) for sample in audio)
return torch.zeros(len(audio), 20, 80), [20] * len(audio), captured[-len(audio):]
wrapper = AutoModel.__new__(AutoModel)
wrapper.model = Probe('asr')
wrapper.vad_model = Probe('vad')
wrapper.spk_model = Probe('speaker')
wrapper.punc_model = None
wrapper.kwargs = {
'device': 'cpu', 'disable_pbar': True, 'return_spk_res': False,
'frontend': SimpleNamespace(fs=16000), 'ncpu': 1,
}
wrapper.vad_kwargs = {'device': 'cpu'}
wrapper.spk_kwargs = {'device': 'cpu'}
wrapper._store_base_configs()
source_rate = 48000
audio = np.sin(2 * np.pi * 440 * np.arange(source_rate // 5) / source_rate).astype(np.float32)
with patch.object(campplus_model, 'extract_feature', capture_features):
wrapper.generate(input=audio, fs=source_rate)
print('sample counts entering speaker feature extraction:', captured)
print('expected: [24000] (1.5 seconds padded at 16 kHz)')
assert captured == [24000], captured
Expected behavior
The speaker model receives the actual rate of its input segments (fs=16000 in this pipeline) and feature extraction sees 24,000 samples. An explicit source rate still reaches VAD/source loading and constructor configuration remains available for the next call.
Error logs
speaker received fs: 48000
sample counts entering speaker feature extraction: [8000]
expected: [24000] (1.5 seconds padded at 16 kHz)
AssertionError: [8000]
The proposed one-line boundary fix produces fs: 16000 and [24000] on the same script. New focused regressions fail 21 cases on the baseline while 8 controls pass; the fixed source passes all 29. Tests cover arrays/explicit PCM, single/batch inputs, constructor/runtime precedence, speaker config conflicts and subsequent calls. No recognition or diarization quality claim is made from these stand-in models.
Environment
- OS: macOS 26.6.2 arm64
- Python: 3.12.12
- FunASR: source 1.4.16, SHA above
- ModelScope: 1.40.1
- PyTorch / torchaudio: 2.10.0 / 2.10.0
- NumPy: 1.26.4; transformers: 4.51.3; tokenizers: 0.21.4
- Install method: editable source in a task-local virtual environment
- Device: CPU; no GPU/CUDA/Docker execution
Audio details
Synthetic mono float32 440 Hz sine, 200 ms, source rate 48 kHz; additional regressions use 8/16 kHz and explicit signed PCM16. Speaker chunking pads it to 1.5 seconds. There is no speech, real speaker identity or background audio.
Prepared with AI assistance; this is a current-source reproduction, not a recovered historical patch.
- 主要语言
- Python
- 星标
- 20.6k
- 派生
- 2.1k
- 平均合并
- 18 小时 3 分钟
- 30 天内合并 PR
- 88
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
modelscope/FunASR 的其他 Issue
-
Bug: blank lines in wav.scp / jsonl file lists crash prepare_data_iterator (IndexError / JSONDecodeError)可能已有人在做 @Lesereingrape 于 4 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 85/100
modelscope/FunASR#3757 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 88/100
modelscope/FunASR#3730 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
modelscope/FunASR#3704 · 1 条评论 ·
维护者通常 1 天内回复
-
没有vllm时使用fun-asr-nano每次都重新加载模型可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭bug needs feedback
难度 2/5 1-3 小时 新手友好度 76/100
modelscope/FunASR#3401 · 2 条评论 ·
维护者通常 1 天内回复
-
Bug: cmvn readers raise IndexError on blank lines and return empty statistics without error可能已有人在做 @Lesereingrape 于 4 天前认领。 未关闭
难度 3/5 1-2 天 新手友好度 35/100
modelscope/FunASR#3759 ·
维护者通常 1 天内回复
查看 modelscope/FunASR 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 85/100
Vector35/community-plugins#376 ·
-
难度 2/5 1-3 小时 新手友好度 68/100
py-econometrics/pyfixest#1883 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
ietf-tools/rfc2html#81 ·
-
难度 1/5 1 小时以内 新手友好度 88/100
mysql/mysql-operator#60 ·
-
Python: Bug: split_plaintext_paragraph / split_markdown_paragraph can return a chunk larger than max_tokens可能已有人在做 @xThreeh 今天认领。 未关闭python triage
难度 2/5 1-3 小时 新手友好度 75/100
microsoft/semantic-kernel#14566 ·
维护者通常 4 天内回复