Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Bug: model.generate(input=bytes) hangs for minutes on raw PCM chunks that trigger audio container false-positive

未關閉
#3,739 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
45/100
Issue 類型
缺陷
描述清晰度
需要釐清
活躍度
活躍
技術堆疊
python, pytorch

研究方向

Read funasr/utils/load_utils.py, especially load_bytes(), _is_audio_container(), and _has_consecutive_mp3_frames(), then trace the input path from funasr/auto/auto_model.py. Build a focused test with 6400-byte raw int16 PCM chunks, verify whether they are misclassified as container audio, and confirm that generate() returns promptly without the reported decode failure.

由索引模型根據 Issue 內容生成。

描述

needs feedback

🐛 Bug Report: model.generate(input=bytes) 在处理流式原始 PCM 时长时间 hang

版本
  • FunASR: 1.4.0
  • Python: 3.10.12
  • PyTorch: 1.12.0
  • OS: Centos7.9
场景

FSMN VAD 模型做流式推理:每 200ms 收到一个 int16 PCM chunk(6400 字节,16kHz mono),调用 AutoModel.generate()。

调用方式
from funasr import AutoModel

model = AutoModel(model="iic/speech_fsmn_vad_zh-cn-16k-common-pytorch", disable_update=True)

# 每 200ms 一帧,pcm_chunk 是原始 int16 PCM bytes,6400 字节(16kHz x 0.2s x 2)
cache = {}
for pcm_chunk in pcm_chunks:
    result = model.generate(
        input=pcm_chunk,      # bytes, 原始 int16 PCM, 6400 字节
        cache=cache,
        chunk_size=200,
        disable_pbar=True,
    )
现象

服务稳定运行数周后,某次处理中 generate() 调用hang 约三分钟不返回。期间该路会话的所有后续音频帧无法处理。三分钟后抛出异常:

Failed to decode container-formatted audio bytes.
Verify that the input is a complete supported audio file and that torchaudio,
soundfile, or ffmpeg is available.

通常 generate() 对一个 200ms chunk 的推理只需 10-20ms。hang 的持续时间远超正常推理时间。

补充信息
  • FunASR 版本 1.4.0(pip show funasr 输出)
  • 输入是原始 int16 PCM bytes,不是 WAV/MP3 等容器格式文件
  • 该错误低频出现(数周一次),未找到稳定复现方式

以下是 AI(Claude)辅助分析的推测,供参考:

AI 推测的可能原因
  1. _is_audio_container 误判:load_utils.py 中 load_bytes() 对所有 bytes 输入调用 _is_audio_container() 进行容器格式检测。200ms int16 PCM chunk(6400 字节)有一定概率命中 MPEG 帧头假阳性(data[0] == 0xFF and (data[1] & 0xE0) == 0xE0),随后 _has_consecutive_mp3_frames() 在随机 PCM 数据中有非零概率通过三层帧头校验。

  2. 误判后 torchaudio 陷入长时间错误恢复:一旦判定为容器格式,load_bytes() 调用 load_audio_text_image_video(BytesIO(input), fs=16000),torchaudio 被喂入一个 6400 字节的"假 MP3 文件",内部错误恢复路径反复尝试重新同步,持续数分钟后才放弃。

(以上推测基于对 FunASR 1.4.1 funasr/utils/load_utils.py 和 funasr/auto/auto_model.py 源码的阅读,尚未通过构造特定 PCM 数据复现验证。)

主要語言
Python
星號
20.5k
分支
2.1k
平均合併
15 小時 21 分鐘
30 天內合併 PR
125

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

modelscope/FunASR 的其他 Issue

查看 modelscope/FunASR 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。