load_pretrained_model deep-copies the whole state dict, doubling peak load memory
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 88/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
調査の方向性
funasr/train_utils/load_pretrained_model.py から始め、torch.load と冗長な deepcopy の周辺にある AutoModel のチェックポイント読み込みパスを追跡します。不要な 2 つ目の state-dict コピーを削除し、その後、指定された大きなチェックポイントで再現を実行して、読み込み時に引き続きすべてのキーが一致したと報告されることを確認します。同時に、ピーク RSS にチェックポイントサイズの 2 つ目のコピーが含まれなくなることも確認します。
索引モデルが issue の本文から書いたものです。
説明
🐛 Bug
load_pretrained_model() deep-copies the entire checkpoint state dict on every model load. The copy is redundant, and it holds a second full copy of the weights in memory for the duration of the load, so peak host memory is roughly doubled by the checkpoint size.
Loading iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt) peaks at 3461 MB RSS. With the copy removed the same load peaks at 2624 MB — a 837MB difference that matches the checkpoint size.
The practical consequence is that a model which fits in memory can still fail to load, and container memory limits have to be set to twice the checkpoint size.
To Reproduce
- Install with:
pip install funasr modelscope kaldi-native-fbank - Run: load any large checkpoint through
AutoModeland sample peak RSS - See: no exception on a roomy host — the symptom is peak RSS; on a constrained host it is an OOM kill
pip install funasr==1.4.16 modelscope kaldi-native-fbank
python - <<'PY'
import resource, time
from funasr import AutoModel
t0 = time.time()
AutoModel(
model="iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online",
device="cpu",
disable_update=True,
)
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{time.time() - t0:.1f}s peak RSS {peak:.0f} MB")
PY
Code sample
funasr/train_utils/load_pretrained_model.py:
ori_state = torch.load(path, map_location=map_location) # fresh, local, never read again
src_state = copy.deepcopy(ori_state) # every tensor copied a second time
src_state = src_state["state_dict"] if "state_dict" in src_state else src_state
src_state = src_state["model_state_dict"] if "model_state_dict" in src_state else src_state
src_state = src_state["model"] if "model" in src_state else src_state
The copy is not needed:
ori_stateis produced inside the function bytorch.load, so it has a single owner and is not shared with the caller.ori_stateis not referenced again after thedeepcopyline.src_stateis only read from that point on. The loop below readssrc_state[k_src].shapeand rebinds references indst_state(dst_state[k] = src_state[k_src]); the data actually reaches the model throughobj.load_state_dict(dst_state, strict=True).- No tensor is mutated in place anywhere on this path, so the deep-copied version and the shared version are equivalent.
Expected behavior
Loading a checkpoint should not require memory for two copies of it. A 220M-parameter model should not need ~2x its own weight size in headroom.
Error logs
No exception. Measured with resource.getrusage(RUSAGE_SELF).ru_maxrss, same machine, same checkpoint, back to back:
funasr 1.4.16 : LOAD 13.3s peak RSS 3461 MB <All keys matched successfully>
same, deep copy removed : LOAD 26.9s peak RSS 2624 MB <All keys matched successfully>
The 837MB difference tracks the 840MB checkpoint, which is the copy being dropped. The wall-clock column is noisy and is not part of the report; the memory difference is the reproducible result. All keys matched successfully on both sides is the correctness check.
Environment
- OS: Linux 6.6.87.2-microsoft-standard-WSL2 (Ubuntu userspace)
- Python version: 3.12.14
- FunASR version: 1.4.16 (latest release;
mainat 41778c4 is identical here) - ModelScope version: 1.40.1
- PyTorch version: 2.14.0+cu126
- Install method:
pip - Device: cpu for the load measurement
- GPU model: NVIDIA GeForce RTX 4060 Laptop GPU
- CUDA version: 12.6
Audio details
Not audio-related. The model used for the measurement is iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt). The overhead scales with checkpoint size, so it should reproduce with any large FunASR checkpoint, including the LLM-ASR models.
- 主要言語
- Python
- スター
- 20.5k
- フォーク
- 2.1k
- 平均マージ
- 12時間 20分
- マージ済み PR(30日)
- 157
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
modelscope/FunASR のほかの issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
modelscope/FunASR#3730 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
modelscope/FunASR#3704 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug needs feedback
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
modelscope/FunASR#3401 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
关于实时模式的vllm解码并发问题オープンneeds triage question
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
modelscope/FunASR#3732 ·
メンテナーはふだん 1 日以内に返信
-
bug needs triage
難易度 4/5 3〜5日 初心者へのやさしさ 30/100
modelscope/FunASR#3727 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
modelscope/FunASR の issue をすべて見る
似ている issue
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedオープンworkflow
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
メンテナーはふだん 1 日以内に返信
-
metadata submission
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
canonical/content-cache-operator#163 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
[submission]オープンsubmission
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
leanprover/lean-eval-submissions#1852 ·
メンテナーはふだん 1 日以内に返信