google/seq2seq

Prepare WMT'17 Datasets

开放

#21 创建于 2017年3月11日

 (3 条评论) (0 个反应) (0 位负责人)Python (1,329 个派生)batch import
datahelp wanted

仓库指标

星标
 (5,587 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

We should prepare datasets for All WMT'17 language pairs. This is also a change to try out google/sentencepiece as a preprocessor.

Each dataset should come in different configurations, i.e. different vocabulary sizes and also have a character-level version.

Together with the raw data files we also need the script that was used for the process.

贡献者指南