Custom video dataset encoding/serialize uses all memory, process killed. How to fix?
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- jupyter-notebook, python
- Lĩnh vực
- data-engineering, machine-learning
Hướng nghiên cứu
Tái hiện mức tăng bộ nhớ bằng các Colab notebooks được liên kết và kiểm tra dgs_corpus.py, VideoFeature tùy chỉnh trong config.py, cùng các đường dẫn split_builder.py và writer.py được dẫn chiếu. Theo dõi nơi các khung hình video được trích xuất, mã hóa và tuần tự hóa, sau đó xác định một thay đổi tập trung với một bài kiểm tra bộ nhớ có thể tái lập và một kết quả rõ ràng về việc giới hạn bộ nhớ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
What I need help with / What I was wondering
I want to load a dataset containing these
without this happening (Colab notebook for replicating)
...How can I edit my dataset loader to use less memory when encoding videos?
Background:
I am trying to load a custom dataset with a Video feature.
When I try to tfds.load() it, or even just download_and_prepare, RAM usage goes up very high and then the process gets killed.
For example this notebook will crash if allowed to run, though with a High-RAM instance it may not.
It seems it is using over 30GB of memory to encode one or two 10 MB videos.
I would like to know how to edit/update this custom dataset so that it will not use so much memory.
What I've tried so far
I did a bunch of debugging and tracing of the problem with memray, etc. See this notebook and this issue for detailed analysis including a copy of the memray report.
Tried various different ideas in the notebook, including loading just a slice, editing buffer size, and switching from .load() to download_and_prepare()
Finally I traced the problem to serializing and encoding steps under the
See this comment, which was allocating many GiB of memory to encode even one 10MB video.
I discovered that even one 10MB video was extracted to over 13k video frames, taking up nearly 5GiB of space. And then the
serializing would take up 14-15 GiB, and the encoding would take another 14-15, and so the process would be killed.
Relevant items:
- The data loader in question, dgs_corpus.py
- The full memray report: memray_output_file.tar.gz
- Encoding path: The dataset uses a custom VideoFeature as well, defined here. The memray showsthat encode_example here ends up allocating 14.5 GiB
- Serialization: The memray shows that the other path that uses memory is serialization: split_builder.py here which calls writer.py's serialization
It would be nice if...
- ...there were more examples of how to efficiently load video datasets, and explanations of why they are more efficient.
- ...there were a way to do this in some sort of streaming fashion that used less memory, e.g. loading in a batch of frames, using a sliding window, etc.
- ...there were some way to set a memory limit, and just have it process more slowly within that limit.
- ...there were a way to separate the download and prepare processes. A download_only option, like
--download_onlyin the CLI - ...there were a warning that the dataset was using a lot of memory in processing, before the OS kills the process.
- ...for saving disk space, a way to encode and serialize videos without extracting thousands of individual frames, ballooning the size from 10MB to multiple GiB. Maybe there is and I just don't know.
- ...it was possible to download only part of a dataset. It's possible to load a slice, but only after download_and_prepare does its whole thing.
- ...more explanation of what serialization and encoding are for, maybe? What are they?
Environment information
I've tested it on Colab and a few other Ubuntu workstations. High-Ram Colab Instances seem to have enough memory to get past this.
- Ngôn ngữ chính
- Python
- Star
- 4.6k
- Fork
- 1.6k
- Merge trung bình
- 5 giờ 47 phút
- Pull request đã merge (30 ngày)
- 2
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của tensorflow/datasets
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
tensorflow/datasets#11212 · 2 bình luận · 1 reaction ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 76/100
tensorflow/datasets#11233 ·
-
dataset request
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 20/100
tensorflow/datasets#11228 ·
-
QM9 download links are brokenĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 55/100
tensorflow/datasets#11210 · 1 bình luận ·
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 1/100
tensorflow/datasets#11196 ·
Tất cả issue của tensorflow/datasets
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
Maintainer thường phản hồi trong vòng 1 ngày
-
https://search.utilibre.orgĐang mởinstance instance add
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
searxng/searx-instances#941 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày