Define tokenizer strategy
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 20/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
调研方向
未指定任何文件、测试或入口点。首先定位 n-gram 和 neural fallback 组件,然后将 tokenizer 选项与结构化的 conventional-commit 语法及罕见术语要求进行比较。完成的标准是:有一份记录完善的策略、关于 special-token 和 vocabulary 的决策、一个可训练的 tokenizer,以及一个通过的 roundtrip 测试。
由索引模型根据 Issue 内容生成。
描述
Summary
Define how conventional commit messages will be tokenized for both n-gram lookup and neural fallback.
Success Criteria
- Tokenization strategy chosen (BPE, WordPiece, character, or hybrid)
- Special tokens defined (, , , etc.)
- Vocabulary size determined
- Tokenizer can be trained on our dataset
- Roundtrip test: tokenize → detokenize preserves meaning
Context
The tokenizer must work well for both:
- N-gram lookup: Needs discrete, interpretable tokens
- Neural fallback: Needs subword handling for rare terms
Considerations
- Conventional commits have structured grammar (type(scope): subject)
- Technical terms may need subword tokenization
- Should we use different tokenizers for different components?
- 主要语言
- Python
- 星标
- 0
- 派生
- 0
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aRustyDev/ccgram 的其他 Issue
-
docs experiment
难度 5/5 一周以上 新手友好度 35/100
-
docs
难度 4/5 3-5 天 新手友好度 55/100
-
infra model
难度 5/5 一周以上 新手友好度 25/100
-
docs experiment
难度 3/5 3-5 天 新手友好度 45/100
-
ablation evaluation
难度 4/5 3-5 天 新手友好度 30/100
相似的 Issue
-
customer-reported
难度 2/5 1-3 小时 新手友好度 68/100
Azure/azure-cli#34150 · 1 条评论 ·
维护者通常 1 天内回复
-
community-request
难度 1/5 1 小时以内 新手友好度 95/100
NVIDIA-NeMo/Curator#2464 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
WeblateOrg/translation-finder#1099 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
trezor/trezor-firmware#7997 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复