Cannot tokenize byte sequences that are not valid UTF-8 due to design flaw
还没有人认领这个 Issue。
评估
调研方向
首先检查 bpe-openai 的 encode 方法及其接受的输入类型。确定库应如何公开任意字节序列,然后为无效 UTF-8 输入添加覆盖测试,并验证所有字节序列都可以进行分词。
由索引模型根据 Issue 内容生成。
描述
Hello,
The BPE algorithm is capable of tokenizing any byte sequence, and LLMs generally accept any sequence of tokens and use token dictionaries that can successfully represent any byte sequence, but the encode method in bpe-openai accepts a type that has to be valid UTF-8. So there are lots of byte sequences, many of which are only 1 byte long, which you cannot tokenize using this library.
- 主要语言
- Rust
- 星标
- 134
- 派生
- 24
- 平均合并
- 16 小时 27 分钟
- 30 天内合并 PR
- 11
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
github/rust-gems 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
-
难度 5/5 一周以上 新手友好度 30/100
-
难度 4/5 3-5 天 新手友好度 38/100
-
难度 5/5 一周以上 新手友好度 25/100
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
stratum-mining/stratum#2404 ·
-
难度 2/5 1-3 小时 新手友好度 85/100
-
难度 2/5 1-3 小时 新手友好度 84/100
Eynzof/Hermes-CN-Desktop#610 ·
-
难度 2/5 1-3 小时 新手友好度 88/100
-
bug team:backend track:services-maintenance
难度 2/5 1-3 小时 新手友好度 78/100
cowprotocol/services#4950 ·