Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Bug: model.generate(input=bytes) hangs for minutes on raw PCM chunks that trigger audio container false-positive

Open
#3,739 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
45/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Active
Tech stack
python, pytorch

Research direction

Read funasr/utils/load_utils.py, especially load_bytes(), _is_audio_container(), and _has_consecutive_mp3_frames(), then trace the input path from funasr/auto/auto_model.py. Build a focused test with 6400-byte raw int16 PCM chunks, verify whether they are misclassified as container audio, and confirm that generate() returns promptly without the reported decode failure.

Written by the indexing model from the issue text.

Description

needs feedback

🐛 Bug Report: model.generate(input=bytes) 在处理流式原始 PCM 时长时间 hang

版本
  • FunASR: 1.4.0
  • Python: 3.10.12
  • PyTorch: 1.12.0
  • OS: Centos7.9
场景

FSMN VAD 模型做流式推理:每 200ms 收到一个 int16 PCM chunk(6400 字节,16kHz mono),调用 AutoModel.generate()。

调用方式
from funasr import AutoModel

model = AutoModel(model="iic/speech_fsmn_vad_zh-cn-16k-common-pytorch", disable_update=True)

# 每 200ms 一帧,pcm_chunk 是原始 int16 PCM bytes,6400 字节(16kHz x 0.2s x 2)
cache = {}
for pcm_chunk in pcm_chunks:
    result = model.generate(
        input=pcm_chunk,      # bytes, 原始 int16 PCM, 6400 字节
        cache=cache,
        chunk_size=200,
        disable_pbar=True,
    )
现象

服务稳定运行数周后,某次处理中 generate() 调用hang 约三分钟不返回。期间该路会话的所有后续音频帧无法处理。三分钟后抛出异常:

Failed to decode container-formatted audio bytes.
Verify that the input is a complete supported audio file and that torchaudio,
soundfile, or ffmpeg is available.

通常 generate() 对一个 200ms chunk 的推理只需 10-20ms。hang 的持续时间远超正常推理时间。

补充信息
  • FunASR 版本 1.4.0(pip show funasr 输出)
  • 输入是原始 int16 PCM bytes,不是 WAV/MP3 等容器格式文件
  • 该错误低频出现(数周一次),未找到稳定复现方式

以下是 AI(Claude)辅助分析的推测,供参考:

AI 推测的可能原因
  1. _is_audio_container 误判:load_utils.py 中 load_bytes() 对所有 bytes 输入调用 _is_audio_container() 进行容器格式检测。200ms int16 PCM chunk(6400 字节)有一定概率命中 MPEG 帧头假阳性(data[0] == 0xFF and (data[1] & 0xE0) == 0xE0),随后 _has_consecutive_mp3_frames() 在随机 PCM 数据中有非零概率通过三层帧头校验。

  2. 误判后 torchaudio 陷入长时间错误恢复:一旦判定为容器格式,load_bytes() 调用 load_audio_text_image_video(BytesIO(input), fs=16000),torchaudio 被喂入一个 6400 字节的"假 MP3 文件",内部错误恢复路径反复尝试重新同步,持续数分钟后才放弃。

(以上推测基于对 FunASR 1.4.1 funasr/utils/load_utils.py 和 funasr/auto/auto_model.py 源码的阅读,尚未通过构造特定 PCM 数据复现验证。)

Dominant language
Python
Stars
20.5k
Forks
2.1k
Avg merge
17h 27m
Merged PRs (30d)
91

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from modelscope/FunASR

All issues in modelscope/FunASR

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.