huggingface/datasets

Return the name of the currently loaded file in the load_dataset function.

开放

#5,806 创建于 2023年4月28日

 (23 条评论) (2 个反应) (1 位负责人)Python (2,496 个派生)batch import
enhancementgood first issue

仓库指标

星标
 (18,313 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

Feature request

Add an optional parameter return_file_name in the load_dataset function. When it is set to True, the function will include the name of the file corresponding to the current line as a feature in the returned output.

Motivation

When training large language models, machine problems may interrupt the training process. In such cases, it is common to load a previously saved checkpoint to resume training. I would like to be able to obtain the names of the previously trained data shards, so that I can skip these parts of the data during continued training to avoid overfitting and redundant training time.

Your contribution

I currently use a dataset in jsonl format, so I am primarily interested in the json format. I suggest adding the file name to the returned table here https://github.com/huggingface/datasets/blob/main/src/datasets/packaged_modules/json/json.py#L92.

贡献者指南