Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

load_pretrained_model deep-copies the whole state dict, doubling peak load memory

Đã đóng Phù hợp với người mới
#3,728 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
88/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python, pytorch
Lĩnh vực
machine-learning

Hướng nghiên cứu

Bắt đầu trong funasr/train_utils/load_pretrained_model.py và lần theo đường dẫn tải checkpoint của AutoModel xung quanh torch.load và deepcopy dư thừa. Xóa bản sao state-dict thứ hai không cần thiết, sau đó chạy lại phép tái hiện với checkpoint lớn được nêu tên và xác nhận rằng quá trình tải vẫn báo cáo tất cả các khóa đều khớp, trong khi RSS cực đại không còn bao gồm một bản sao thứ hai có kích thước bằng checkpoint.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug needs triage

🐛 Bug

load_pretrained_model() deep-copies the entire checkpoint state dict on every model load. The copy is redundant, and it holds a second full copy of the weights in memory for the duration of the load, so peak host memory is roughly doubled by the checkpoint size.

Loading iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt) peaks at 3461 MB RSS. With the copy removed the same load peaks at 2624 MB — a 837MB difference that matches the checkpoint size.

The practical consequence is that a model which fits in memory can still fail to load, and container memory limits have to be set to twice the checkpoint size.

To Reproduce

  1. Install with: pip install funasr modelscope kaldi-native-fbank
  2. Run: load any large checkpoint through AutoModel and sample peak RSS
  3. See: no exception on a roomy host — the symptom is peak RSS; on a constrained host it is an OOM kill
pip install funasr==1.4.16 modelscope kaldi-native-fbank

python - <<'PY'
import resource, time
from funasr import AutoModel

t0 = time.time()
AutoModel(
    model="iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online",
    device="cpu",
    disable_update=True,
)
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{time.time() - t0:.1f}s   peak RSS {peak:.0f} MB")
PY

Code sample

funasr/train_utils/load_pretrained_model.py:

ori_state = torch.load(path, map_location=map_location)   # fresh, local, never read again

src_state = copy.deepcopy(ori_state)                     # every tensor copied a second time
src_state = src_state["state_dict"] if "state_dict" in src_state else src_state
src_state = src_state["model_state_dict"] if "model_state_dict" in src_state else src_state
src_state = src_state["model"] if "model" in src_state else src_state

The copy is not needed:

  • ori_state is produced inside the function by torch.load, so it has a single owner and is not shared with the caller.
  • ori_state is not referenced again after the deepcopy line.
  • src_state is only read from that point on. The loop below reads src_state[k_src].shape and rebinds references in dst_state (dst_state[k] = src_state[k_src]); the data actually reaches the model through obj.load_state_dict(dst_state, strict=True).
  • No tensor is mutated in place anywhere on this path, so the deep-copied version and the shared version are equivalent.

Expected behavior

Loading a checkpoint should not require memory for two copies of it. A 220M-parameter model should not need ~2x its own weight size in headroom.

Error logs

No exception. Measured with resource.getrusage(RUSAGE_SELF).ru_maxrss, same machine, same checkpoint, back to back:

funasr 1.4.16              : LOAD 13.3s   peak RSS 3461 MB   <All keys matched successfully>
same, deep copy removed    : LOAD 26.9s   peak RSS 2624 MB   <All keys matched successfully>

The 837MB difference tracks the 840MB checkpoint, which is the copy being dropped. The wall-clock column is noisy and is not part of the report; the memory difference is the reproducible result. All keys matched successfully on both sides is the correctness check.

Environment

  • OS: Linux 6.6.87.2-microsoft-standard-WSL2 (Ubuntu userspace)
  • Python version: 3.12.14
  • FunASR version: 1.4.16 (latest release; main at 41778c4 is identical here)
  • ModelScope version: 1.40.1
  • PyTorch version: 2.14.0+cu126
  • Install method: pip
  • Device: cpu for the load measurement
  • GPU model: NVIDIA GeForce RTX 4060 Laptop GPU
  • CUDA version: 12.6

Audio details

Not audio-related. The model used for the measurement is iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online (220M params, 956 tensors, 840MB model.pt). The overhead scales with checkpoint size, so it should reproduce with any large FunASR checkpoint, including the LLM-ASR models.

Ngôn ngữ chính
Python
Star
20.5k
Fork
2.1k
Merge trung bình
12 giờ 20 phút
Pull request đã merge (30 ngày)
157

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của modelscope/FunASR

Tất cả issue của modelscope/FunASR

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.