Cannot tokenize byte sequences that are not valid UTF-8 due to design flaw

未关闭
#51 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
需要澄清
活跃度
停滞
技术栈
rust
领域
tooling

调研方向

首先检查 bpe-openai 的 encode 方法及其接受的输入类型。确定库应如何公开任意字节序列,然后为无效 UTF-8 输入添加覆盖测试,并验证所有字节序列都可以进行分词。

由索引模型根据 Issue 内容生成。

描述

Hello,

The BPE algorithm is capable of tokenizing any byte sequence, and LLMs generally accept any sequence of tokens and use token dictionaries that can successfully represent any byte sequence, but the encode method in bpe-openai accepts a type that has to be valid UTF-8. So there are lots of byte sequences, many of which are only 1 byte long, which you cannot tokenize using this library.

主要语言
Rust
星标
134
派生
24
平均合并
16 小时 27 分钟
30 天内合并 PR
11

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

github/rust-gems 的其他 Issue

查看 github/rust-gems 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。