Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Feature]: HF tokenizer.json parity for BPE encoders — array-form merges, normalizer, non-special added tokens, Metaspace BPE

オープン
#1,321 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
kotlin, python
領域
ai, backend

調査の方向性

Start by reading TokenizerFactory.kt, QwenByteLevelBpeTokenizer.kt, and the existing TokenizerFactoryDispatchTest; compare the documented merge parsing and added token handling with Hugging Face tokenizers 0.23.3. The issue divides the work into separate lanes, with NFC strategy and Metaspace BPE requiring research and design. Done means both published Laya tokenizer files load without rewriting and match the reference fixtures, with the stated vocabulary size and decode checks.

索引モデルが issue の本文から書いたものです。

説明

darc enhancement tracking

🧠 D: DEFINE — Problem & Opportunity

sk.ainet.io.tokenizer.TokenizerFactory.fromTokenizerJson (skainet-io-core 0.57.0) cannot load the tokenizer.json of any current Hugging Face encoder checkpoint, and when the file is patched so that it loads, the ids it produces differ from the tokenizers reference. Measured on the two tokenizer files shipped with the Laya checkpoints — ModernBERT-large (byte-level BPE, NFC normalizer) and mmBERT-base (SentencePiece-style BPE, Metaspace pre-tokenizer, byte fallback):

tokenizer.json As published With model.merges rewritten to the legacy "a b" string form
ModernBERT (tokenizer/tokenizer.json) load fails: IllegalArgumentException: Element class kotlinx.serialization.json.JsonArray … is not a JsonPrimitive QwenByteLevelBpeTokenizer, 3 of 5 sample texts identical to the reference
mmBERT (multilingual/tokenizer/tokenizer.json) same load failure QwenByteLevelBpeTokenizer, 0 of 5 identical (byte-level pre-tokenization applied to a ▁ vocabulary)

Four independent gaps, in increasing size:

  1. Array form of model.merges. the Hugging Face tokenizers Python library (PyPI; not a SKaiNET version) writes merges since its 0.20 release as [["a","b"], …] (verified with 0.20.3 and with the current 0.23.3: both write the array form and still read the string form). The builder at QwenByteLevelBpeTokenizer.kt#L243-L255 only accepts "a b" strings, so every freshly exported tokenizer fails to load.
  2. Non-special added_tokens are dropped. The special: false entries are filtered out at QwenByteLevelBpeTokenizer.kt#L257-L267 and TokenizerFactory.kt#L155-L169. The reference matches every added token atomically before pre-tokenization; special only controls skip_special_tokens on decode. ModernBERT has 109 such tokens (whitespace runs, |||IP_ADDRESS|||): " leading" → reference [50275, 16378], engine [245, 4283, …]. vocabSize also reports 50280 instead of 50368.
  3. The normalizer block is ignored. ModernBERT declares {"type": "NFC"}; a decomposed e + U+0301 is encoded as three byte tokens instead of the composed é token (reference […, 3974], engine […, 299, 136, 212]). Only SentencePieceTokenizer looks at the block, and only to sniff Prepend (SentencePieceTokenizer.kt#L382-L391).
  4. Metaspace BPE with byte fallback. model.type: "BPE" + pre_tokenizer: Metaspace + byte_fallback: true is the Llama/Gemma/mmBERT family. It is routed to the byte-level class, which applies the GPT-2 regex and the byte→unicode table to a ▁ vocabulary. This is the encode-side twin of #912 (which reports the decode-side crash on the same files).

Who benefits: anyone loading a SafeTensors checkpoint with its tokenizer.json — encoder models (ModernBERT, mmBERT, BERT-class rerankers) and every decoder exported with a current tokenizers. The downstream SKaiNET-transformers factory delegates model.type: BPE straight here (UpstreamTokenizerAdapter), so the fix lands in one place.

Summary (1–2 sentences):

Make TokenizerFactory.fromTokenizerJson reproduce the Hugging Face tokenizers ids for BPE tokenizer.json files: accept array-form merges, match all added tokens, apply the declared normalizer, and dispatch Metaspace BPE to a SentencePiece-style BPE encoder instead of the byte-level one.


🔍 A: ASSESS — Feasibility & Impact

✔️ Feasibility
  • Gaps 1 and 2 are local, few-line changes with existing synthetic tests to extend (TokenizerFactoryDispatchTest). Both are good first issues.
  • Gap 3 needs an NFC/NFKC implementation reachable from commonMain (JVM has java.text.Normalizer; Kotlin/Native, JS and Wasm do not ship one). Whether to go expect/actual, table-driven, or JVM-first-with-passthrough is the one design decision in this feature → research lane first.
  • Gap 4 is the largest: the merge-rank BPE core in QwenByteLevelBpeTokenizer is correct, but its alphabet mapping (bytes → unicode, GPT-2 regex) is hard-wired. Splitting the BPE core from the pre-tokenizer/alphabet step lets one core serve ByteLevel and Metaspace. SentencePieceTokenizer is not a substitute: it implements llama.cpp score-priority merging over a Unigram vocabulary and expects [token, score] pairs.
✔️ Expected Impact
  • Every tokenizer.json exported by tokenizers ≥ 0.20 loads again (gap 1 alone unblocks this).
  • Encoder-class checkpoints become loadable without a per-model tokenizer fork.
  • Closes the encode half of #912; the decode half (decode throwing on ▁, bos/eos from added_tokens) is finished by the same lane.
✔️ Risks / Constraints
  • Changing vocabSize to include added tokens changes a public value; callers that size embedding tables from it get the correct number afterwards, but it is a behaviour change — list it in the changelog.
  • Added-token matching must respect lstrip/rstrip/single_word/normalized flags; the Laya files use only normalized: true on the non-special entries. Anything else → separate issue, do not block.
  • NFC on non-JVM targets: if the research lane finds no cheap table, ship JVM actual + a documented passthrough elsewhere rather than blocking the whole feature.
✔️ Dependencies
  • skainet-io-core only (sk.ainet.io.tokenizer). No new third-party library unless the NFC research recommends one (it should not).
  • Reference oracle: Hugging Face tokenizers (Python), pinned per golden file.

📚 R: RESEARCH — What Must Be Understood First?

Research Tasks
  • Survey which normalizer / pre_tokenizer / decoder combinations appear in the tokenizer.json of popular public checkpoints (Lane 1).
  • Decide the NFC strategy for commonMain (Lane 1 output, consumed by the normalizer lane).
  • Pin the exact matching semantics of added_tokens flags in tokenizers (AddedVocabulary), with citations.
Open Questions
  • Should vocabSize include added tokens whose id is above model.vocab? (Reference: yes — get_vocab_size(with_added_tokens=True).)
  • Does the ignore_merges: true flag (Llama 3 files) need handling in the same pass as gap 4?

🛠️ C: CONTRIBUTE — Implementation Plan

Development Tasks
  • Lane 1 · numerics — normalizer / pre-tokenizer survey and NFC recommendation
  • Lane 2 · kotlin-core — array-form model.merges (good first issue)
  • Lane 2 · kotlin-core — atomic non-special added_tokens + vocabSize (good first issue)
  • Lane 2 · kotlin-core — apply the normalizer block (blocked by Lane 1)
  • Lane 2 · kotlin-core — Metaspace BPE with byte fallback (encode side of #912)
  • Lane 4 · numerics — golden encode fixtures from tokenizers for the two Laya files (good first issue)
  • Lane 4 · kotlin-core — fixture-gated jvmTest replaying the goldens (good first issue, blocked by the golden lane)
  • Lane 5 · docs — tokenizer.json support matrix (good first issue)
  • Lane 6 · review — DARC review
Acceptance Criteria
  • Both Laya tokenizer.json files load as published (no merge rewriting) and encode the fixture texts byte-identically to the current tokenizers release (0.23.3 at filing time) (add_special_tokens=False).
  • vocabSize equals get_vocab_size(with_added_tokens=True).
  • The #912 decode fixture passes on JVM and on linuxX64Test.
  • Existing Qwen fixture tests unchanged.
  • No regressions elsewhere; reviewed by someone who did not implement the lanes.

💬 Additional Notes

Fixture texts (shared by every lane; reference ids are produced with tokenizers 0.23.3 (identical under 0.20.3), Tokenizer.from_file(path).encode(text, add_special_tokens=False).ids):

Bitte die doppelte Rechnung prüfen.
Which department handles this request?
Straße, Übergrößen und Ärger – naïve café façade 😀👍🏽 é      ← the final "é" is e + U+0301 (decomposed)
false: No duplicate is mentioned
   leading spaces and 12345 numbers, hyphen-ated words don't

Reference files (Apache-2.0, public): ModernBERT tokenizer/tokenizer.json — normalizer: NFC, pre_tokenizer: ByteLevel, 50,280 vocab + 116 added tokens (7 special); mmBERT multilingual/tokenizer/tokenizer.json — normalizer: Replace(" "→"▁"), pre_tokenizer: Metaspace(prepend_scheme: always), byte_fallback: true, 256,000 vocab, 580,604 merges, 255 <0xNN> tokens.

Lane breakdown

Lane Applies? Why
0 Design (SKEEP) no Additive behaviour behind the existing Tokenizer interface and factory; no API shape change. If the Metaspace lane ends up needing a new public class, flag it there.
1 Numerics / Research yes Normalizer survey + NFC strategy; Python only.
2 Kotlin core yes Four tasks, two of them first issues.
3 Platform verification folded in allTests and linuxX64Test already run skainet-io-core in CI; the fixture lane asserts on both.
4 Ground truth / CI yes Python goldens + Kotlin replay test.
5 Docs yes Support matrix page; the engine docs have no tokenizer page today.
6 Review yes

Related: #912 (decode side, native tests), #858 (legacy files without model.type), #463 (original byte-level BPE fix). Downstream consumer: SKaiNET-transformers TokenizerFactory.fromTokenizerJsonString, which delegates model.type: BPE here.

主要言語
Kotlin
スター
52
フォーク
15
平均マージ
1日 15時間
マージ済み PR(30日)
36

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

SKaiNET-developers/SKaiNET のほかの issue

SKaiNET-developers/SKaiNET の issue をすべて見る

似ている issue

Kotlin の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。