Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Custom video dataset encoding/serialize uses all memory, process killed. How to fix?

オープン
#5,499 コメント 10 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
静か
技術スタック
jupyter-notebook, python

調査の方向性

リンクされた Colab notebooks でメモリ増加を再現し、dgs_corpus.py、config.py のカスタム VideoFeature、引用されている split_builder.py と writer.py のパスを調査します。動画フレームがどこで抽出、エンコード、シリアライズされるのかを追跡し、再現可能なメモリテストと、メモリ使用量が上限内に収まることが明確な結果を伴う、焦点を絞った変更を定義します。

索引モデルが issue の本文から書いたものです。

説明

help

What I need help with / What I was wondering

I want to load a dataset containing these
image
without this happening (Colab notebook for replicating)
image

...How can I edit my dataset loader to use less memory when encoding videos?

Background:
I am trying to load a custom dataset with a Video feature.
When I try to tfds.load() it, or even just download_and_prepare, RAM usage goes up very high and then the process gets killed.
For example this notebook will crash if allowed to run, though with a High-RAM instance it may not.
It seems it is using over 30GB of memory to encode one or two 10 MB videos.
I would like to know how to edit/update this custom dataset so that it will not use so much memory.

What I've tried so far
image

I did a bunch of debugging and tracing of the problem with memray, etc. See this notebook and this issue for detailed analysis including a copy of the memray report.

Tried various different ideas in the notebook, including loading just a slice, editing buffer size, and switching from .load() to download_and_prepare()

Finally I traced the problem to serializing and encoding steps under the
See this comment, which was allocating many GiB of memory to encode even one 10MB video.

I discovered that even one 10MB video was extracted to over 13k video frames, taking up nearly 5GiB of space. And then the
serializing would take up 14-15 GiB, and the encoding would take another 14-15, and so the process would be killed.

Relevant items:

It would be nice if...

  • ...there were more examples of how to efficiently load video datasets, and explanations of why they are more efficient.
  • ...there were a way to do this in some sort of streaming fashion that used less memory, e.g. loading in a batch of frames, using a sliding window, etc.
  • ...there were some way to set a memory limit, and just have it process more slowly within that limit.
  • ...there were a way to separate the download and prepare processes. A download_only option, like --download_only in the CLI
  • ...there were a warning that the dataset was using a lot of memory in processing, before the OS kills the process.
  • ...for saving disk space, a way to encode and serialize videos without extracting thousands of individual frames, ballooning the size from 10MB to multiple GiB. Maybe there is and I just don't know.
  • ...it was possible to download only part of a dataset. It's possible to load a slice, but only after download_and_prepare does its whole thing.
  • ...more explanation of what serialization and encoding are for, maybe? What are they?

Environment information
I've tested it on Colab and a few other Ubuntu workstations. High-Ram Colab Instances seem to have enough memory to get past this.

主要言語
Python
スター
4.6k
フォーク
1.6k
平均マージ
5時間 47分
マージ済み PR(30日)
2

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

tensorflow/datasets のほかの issue

tensorflow/datasets の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。