facebookresearch/metaseq

Unify tokenizers

开放

#308 创建于 2022年8月20日

 (1 条评论) (1 个反应) (1 位负责人)Python (701 个派生)batch import
better-engenhancementgood first issue

仓库指标

星标
 (6,195 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

🚀 Feature Request

With #305, we now have two ways to specify a tokenizer: with the GPT2 tokenizer (provided as two files), and with the universal HF format (specified as one file). These are in two separate code paths, but they don't need to be: we could (manually) merge the two GPT2 files into the universal HF format and switch to only that, and we should.

The resulting code would be cleaner, but catching all the other places the old method is used (e.g. API and old sweeps) needs thorough review.

贡献者指南