Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Lane 2 · kotlin-core] tokenizer.json parity — Metaspace BPE with byte fallback (encode side of #912)

Abierto
#1,326 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
1-2 días
Aptitud para principiantes
55/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
kotlin

Línea de trabajo

Start by reading TokenizerFactory.fromTokenizerJson, QwenByteLevelBpeTokenizer, and the existing bpeMerge implementation; the issue describes the needed Metaspace behavior and dispatch rules. Reuse the byte-token lookup and decode logic in SentencePieceTokenizer, then add the synthetic dispatch test and Moonshine decode fixture in the named tests. Done means the specified fixtures pass on JVM and linuxX64, Qwen tests remain unchanged, and results are reported on #1321.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

coding size:m skill:kotlin-core sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 2 · Kotlin core
Skill needed: Kotlin, and you should understand BPE merging (the existing bpeMerge in QwenByteLevelBpeTokenizer is 20 lines and is the algorithm). Familiarity with the SentencePiece ▁ convention helps.
Size: m (1–2 days)
Blocked by: the array-merges lane (the mmBERT file does not load without it). Independent of the normalizer lane, but the mmBERT file also declares Replace(" " → "▁"), so run the fixture after both land.

What to do

  1. Today TokenizerFactory.fromTokenizerJson sends every model.type: "BPE" file to QwenByteLevelBpeTokenizer, which (a) splits with the GPT-2 regex, (b) maps UTF-8 bytes to the GPT-2 unicode alphabet, (c) merges by rank. For pre_tokenizer: Metaspace files (Llama, Gemma HF exports, mmBERT) steps (a) and (b) are wrong: the alphabet is plain code points with ▁ for space, and characters missing from the vocab fall back to <0xNN> byte tokens when model.byte_fallback is true. Every word boundary in the mmBERT file currently becomes id 241753.
  2. Separate the merge-rank BPE core from the alphabet step. Suggested shape: keep QwenByteLevelBpeTokenizer as the ByteLevel configuration and add a MetaspaceBpeTokenizer (or a strategy parameter) that: replaces " " with ▁; honours prepend_scheme (always → prepend ▁ to the text; first → only to the first segment; never); splits on ▁ boundaries when split: true; runs the same rank-based bpeMerge; and emits <0xNN> ids for any symbol not in the vocab when byte_fallback is true (fuse_unk otherwise). SentencePieceTokenizer already has the <0xNN> lookup and decode — reuse, don't copy.
  3. Dispatch in TokenizerFactory.fromTokenizerJson on pre_tokenizer.type (Metaspace → new class, ByteLevel or absent → existing class, Sequence → inspect members).
  4. Decode: ▁ → space, <0xNN> runs → bytes, matching the file's decoder: Sequence[Replace, ByteFallback, Fuse]. This is the crash reported in #912 (decode: token '▁Ever' contains non-byte-level char U+2581); add that issue's fixture (decode([18274, 1898, 29973, 2]) → "Ever tried?" on the Moonshine v2 file) as a test.
  5. Tests: synthetic Metaspace file in TokenizerFactoryDispatchTest (3-token vocab, one merge, one <0x..> fallback); then the mmBERT fixture via the Lane 4 goldens. Run ./gradlew :skainet-io:skainet-io-core:jvmTest :skainet-io:skainet-io-core:linuxX64Test.

Acceptance

  • mmBERT file (multilingual/tokenizer/tokenizer.json) encodes all five #1321 fixture texts identically to the reference; e.g. "Which department handles this request?" → [12236, 9888, 26446, 736, 3853, 235336]
  • #912 decode fixture passes on JVM and linuxX64Test
  • bosTokenId/eosTokenId come from added_tokens (<bos>=2, <eos>=1 in the mmBERT file)
  • Qwen fixture tests (QwenByteLevelBpeTokenizerFixtureTest) unchanged
  • Result reported back on #1321

Notes

If the refactor forces a new public class or changes the Tokenizer interface, stop and say so on #1321 — that trips the SKEEP trigger ("public Kotlin API") and the maintainers decide whether a short SKEEP is needed before the PR. Llama 3 files additionally set ignore_merges: true (vocab lookup of the whole word before merging); note whether you handled it, a follow-up issue is fine.

Lenguaje dominante
Kotlin
Estrellas
52
Forks
15
Merge medio
1 d 15 h
PR fusionados (30 d)
36

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de SKaiNET-developers/SKaiNET

Todos los issues de SKaiNET-developers/SKaiNET

Issues similares

Más issues de Kotlin

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.