load_pretrained_model deep-copies the whole state dict, doubling peak load memory
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 88/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- machine-learning
Hướng nghiên cứu
Bắt đầu trong funasr/train_utils/load_pretrained_model.py và lần theo đường dẫn tải checkpoint của AutoModel xung quanh torch.load và deepcopy dư thừa. Xóa bản sao state-dict thứ hai không cần thiết, sau đó chạy lại phép tái hiện với checkpoint lớn được nêu tên và xác nhận rằng quá trình tải vẫn báo cáo tất cả các khóa đều khớp, trong khi RSS cực đại không còn bao gồm một bản sao thứ hai có kích thước bằng checkpoint.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
🐛 Bug
load_pretrained_model() deep-copies the entire checkpoint state dict on every model load. The copy is redundant, and it holds a second full copy of the weights in memory for the duration of the load, so peak host memory is roughly doubled by the checkpoint size.
Loading iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt) peaks at 3461 MB RSS. With the copy removed the same load peaks at 2624 MB — a 837MB difference that matches the checkpoint size.
The practical consequence is that a model which fits in memory can still fail to load, and container memory limits have to be set to twice the checkpoint size.
To Reproduce
- Install with:
pip install funasr modelscope kaldi-native-fbank - Run: load any large checkpoint through
AutoModeland sample peak RSS - See: no exception on a roomy host — the symptom is peak RSS; on a constrained host it is an OOM kill
pip install funasr==1.4.16 modelscope kaldi-native-fbank
python - <<'PY'
import resource, time
from funasr import AutoModel
t0 = time.time()
AutoModel(
model="iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online",
device="cpu",
disable_update=True,
)
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{time.time() - t0:.1f}s peak RSS {peak:.0f} MB")
PY
Code sample
funasr/train_utils/load_pretrained_model.py:
ori_state = torch.load(path, map_location=map_location) # fresh, local, never read again
src_state = copy.deepcopy(ori_state) # every tensor copied a second time
src_state = src_state["state_dict"] if "state_dict" in src_state else src_state
src_state = src_state["model_state_dict"] if "model_state_dict" in src_state else src_state
src_state = src_state["model"] if "model" in src_state else src_state
The copy is not needed:
ori_stateis produced inside the function bytorch.load, so it has a single owner and is not shared with the caller.ori_stateis not referenced again after thedeepcopyline.src_stateis only read from that point on. The loop below readssrc_state[k_src].shapeand rebinds references indst_state(dst_state[k] = src_state[k_src]); the data actually reaches the model throughobj.load_state_dict(dst_state, strict=True).- No tensor is mutated in place anywhere on this path, so the deep-copied version and the shared version are equivalent.
Expected behavior
Loading a checkpoint should not require memory for two copies of it. A 220M-parameter model should not need ~2x its own weight size in headroom.
Error logs
No exception. Measured with resource.getrusage(RUSAGE_SELF).ru_maxrss, same machine, same checkpoint, back to back:
funasr 1.4.16 : LOAD 13.3s peak RSS 3461 MB <All keys matched successfully>
same, deep copy removed : LOAD 26.9s peak RSS 2624 MB <All keys matched successfully>
The 837MB difference tracks the 840MB checkpoint, which is the copy being dropped. The wall-clock column is noisy and is not part of the report; the memory difference is the reproducible result. All keys matched successfully on both sides is the correctness check.
Environment
- OS: Linux 6.6.87.2-microsoft-standard-WSL2 (Ubuntu userspace)
- Python version: 3.12.14
- FunASR version: 1.4.16 (latest release;
mainat 41778c4 is identical here) - ModelScope version: 1.40.1
- PyTorch version: 2.14.0+cu126
- Install method:
pip - Device: cpu for the load measurement
- GPU model: NVIDIA GeForce RTX 4060 Laptop GPU
- CUDA version: 12.6
Audio details
Not audio-related. The model used for the measurement is iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt). The overhead scales with checkpoint size, so it should reproduce with any large FunASR checkpoint, including the LLM-ASR models.
- Ngôn ngữ chính
- Python
- Star
- 20.5k
- Fork
- 2.1k
- Merge trung bình
- 12 giờ 20 phút
- Pull request đã merge (30 ngày)
- 157
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của modelscope/FunASR
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
modelscope/FunASR#3730 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
modelscope/FunASR#3704 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
没有vllm时使用fun-asr-nano每次都重新加载模型Đang mởbug needs feedback
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
modelscope/FunASR#3401 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
关于实时模式的vllm解码并发问题Đang mởneeds triage question
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
modelscope/FunASR#3732 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
【暴露的问题比断句严重】实时语音,切分不够彻底,怎么解决Đang mởbug needs triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 30/100
modelscope/FunASR#3727 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của modelscope/FunASR
Issue tương tự
-
ACK_WAITING HELP_WANTED UPDATE_CS
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
OWASP/CheatSheetSeries#2458 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 90/100
BasedHardware/omi#19711 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Qwen3_5MoeModel no longer returns router_logits, breaking aux loss with output_router_logits=TrueĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
huggingface/transformers#49172 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
vllm-project/vllm-metal#885 ·
Maintainer thường phản hồi trong vòng 1 ngày