google/seq2seq

Prepare WMT'17 Datasets

開放

#21 建立於 2017年3月11日

 (3 則留言) (0 個反應) (0 位負責人)Python (1,329 個分叉)batch import
datahelp wanted

倉庫指標

星標
 (5,587 顆星)
PR 合併指標
 (PR 指標待抓取)

描述

We should prepare datasets for All WMT'17 language pairs. This is also a change to try out google/sentencepiece as a preprocessor.

Each dataset should come in different configurations, i.e. different vocabulary sizes and also have a character-level version.

Together with the raw data files we also need the script that was used for the process.

貢獻者指南