Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Lane 2 · kotlin-core] tokenizer.json parity — Metaspace BPE with byte fallback (encode side of #912)

Open
#1,326 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
1-2 days
Newbie friendliness
55/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
kotlin

Research direction

Start by reading TokenizerFactory.fromTokenizerJson, QwenByteLevelBpeTokenizer, and the existing bpeMerge implementation; the issue describes the needed Metaspace behavior and dispatch rules. Reuse the byte-token lookup and decode logic in SentencePieceTokenizer, then add the synthetic dispatch test and Moonshine decode fixture in the named tests. Done means the specified fixtures pass on JVM and linuxX64, Qwen tests remain unchanged, and results are reported on #1321.

Written by the indexing model from the issue text.

Description

coding size:m skill:kotlin-core sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 2 · Kotlin core
Skill needed: Kotlin, and you should understand BPE merging (the existing bpeMerge in QwenByteLevelBpeTokenizer is 20 lines and is the algorithm). Familiarity with the SentencePiece ▁ convention helps.
Size: m (1–2 days)
Blocked by: the array-merges lane (the mmBERT file does not load without it). Independent of the normalizer lane, but the mmBERT file also declares Replace(" " → "▁"), so run the fixture after both land.

What to do

  1. Today TokenizerFactory.fromTokenizerJson sends every model.type: "BPE" file to QwenByteLevelBpeTokenizer, which (a) splits with the GPT-2 regex, (b) maps UTF-8 bytes to the GPT-2 unicode alphabet, (c) merges by rank. For pre_tokenizer: Metaspace files (Llama, Gemma HF exports, mmBERT) steps (a) and (b) are wrong: the alphabet is plain code points with ▁ for space, and characters missing from the vocab fall back to <0xNN> byte tokens when model.byte_fallback is true. Every word boundary in the mmBERT file currently becomes id 241753.
  2. Separate the merge-rank BPE core from the alphabet step. Suggested shape: keep QwenByteLevelBpeTokenizer as the ByteLevel configuration and add a MetaspaceBpeTokenizer (or a strategy parameter) that: replaces " " with ▁; honours prepend_scheme (always → prepend ▁ to the text; first → only to the first segment; never); splits on ▁ boundaries when split: true; runs the same rank-based bpeMerge; and emits <0xNN> ids for any symbol not in the vocab when byte_fallback is true (fuse_unk otherwise). SentencePieceTokenizer already has the <0xNN> lookup and decode — reuse, don't copy.
  3. Dispatch in TokenizerFactory.fromTokenizerJson on pre_tokenizer.type (Metaspace → new class, ByteLevel or absent → existing class, Sequence → inspect members).
  4. Decode: ▁ → space, <0xNN> runs → bytes, matching the file's decoder: Sequence[Replace, ByteFallback, Fuse]. This is the crash reported in #912 (decode: token '▁Ever' contains non-byte-level char U+2581); add that issue's fixture (decode([18274, 1898, 29973, 2]) → "Ever tried?" on the Moonshine v2 file) as a test.
  5. Tests: synthetic Metaspace file in TokenizerFactoryDispatchTest (3-token vocab, one merge, one <0x..> fallback); then the mmBERT fixture via the Lane 4 goldens. Run ./gradlew :skainet-io:skainet-io-core:jvmTest :skainet-io:skainet-io-core:linuxX64Test.

Acceptance

  • mmBERT file (multilingual/tokenizer/tokenizer.json) encodes all five #1321 fixture texts identically to the reference; e.g. "Which department handles this request?" → [12236, 9888, 26446, 736, 3853, 235336]
  • #912 decode fixture passes on JVM and linuxX64Test
  • bosTokenId/eosTokenId come from added_tokens (<bos>=2, <eos>=1 in the mmBERT file)
  • Qwen fixture tests (QwenByteLevelBpeTokenizerFixtureTest) unchanged
  • Result reported back on #1321

Notes

If the refactor forces a new public class or changes the Tokenizer interface, stop and say so on #1321 — that trips the SKEEP trigger ("public Kotlin API") and the maintainers decide whether a short SKEEP is needed before the PR. Llama 3 files additionally set ignore_merges: true (vocab lookup of the whole word before merging); note whether you handled it, a follow-up issue is fine.

Dominant language
Kotlin
Stars
52
Forks
15
Avg merge
1d 15h
Merged PRs (30d)
36

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SKaiNET-developers/SKaiNET

All issues in SKaiNET-developers/SKaiNET

Similar issues

More Kotlin issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.