google/seq2seq

Prepare WMT'17 Datasets

オープン

#21 opened on 2017/03/11

 (3 件のコメント) (0 件のリアクション) (0 人の担当者)Python (1,329 件のフォーク)batch import
datahelp wanted

Repository metrics

Stars
 (5,587 個のスター)
PR merge metrics
 (PR metrics pending)

説明

We should prepare datasets for All WMT'17 language pairs. This is also a change to try out google/sentencepiece as a preprocessor.

Each dataset should come in different configurations, i.e. different vocabulary sizes and also have a character-level version.

Together with the raw data files we also need the script that was used for the process.

コントリビューターガイド