[Feature]: HF tokenizer.json parity for BPE encoders — array-form merges, normalizer, non-special added tokens, Metaspace BPE
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
Línea de trabajo
Start by reading TokenizerFactory.kt, QwenByteLevelBpeTokenizer.kt, and the existing TokenizerFactoryDispatchTest; compare the documented merge parsing and added token handling with Hugging Face tokenizers 0.23.3. The issue divides the work into separate lanes, with NFC strategy and Metaspace BPE requiring research and design. Done means both published Laya tokenizer files load without rewriting and match the reference fixtures, with the stated vocabulary size and decode checks.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
🧠 D: DEFINE — Problem & Opportunity
sk.ainet.io.tokenizer.TokenizerFactory.fromTokenizerJson (skainet-io-core 0.57.0) cannot load the tokenizer.json of any current Hugging Face encoder checkpoint, and when the file is patched so that it loads, the ids it produces differ from the tokenizers reference. Measured on the two tokenizer files shipped with the Laya checkpoints — ModernBERT-large (byte-level BPE, NFC normalizer) and mmBERT-base (SentencePiece-style BPE, Metaspace pre-tokenizer, byte fallback):
tokenizer.json |
As published | With model.merges rewritten to the legacy "a b" string form |
|---|---|---|
| ModernBERT (tokenizer/tokenizer.json) | load fails: IllegalArgumentException: Element class kotlinx.serialization.json.JsonArray … is not a JsonPrimitive |
QwenByteLevelBpeTokenizer, 3 of 5 sample texts identical to the reference |
| mmBERT (multilingual/tokenizer/tokenizer.json) | same load failure | QwenByteLevelBpeTokenizer, 0 of 5 identical (byte-level pre-tokenization applied to a ▁ vocabulary) |
Four independent gaps, in increasing size:
- Array form of
model.merges. the Hugging FacetokenizersPython library (PyPI; not a SKaiNET version) writes merges since its 0.20 release as[["a","b"], …](verified with 0.20.3 and with the current 0.23.3: both write the array form and still read the string form). The builder at QwenByteLevelBpeTokenizer.kt#L243-L255 only accepts"a b"strings, so every freshly exported tokenizer fails to load. - Non-special
added_tokensare dropped. Thespecial: falseentries are filtered out at QwenByteLevelBpeTokenizer.kt#L257-L267 and TokenizerFactory.kt#L155-L169. The reference matches every added token atomically before pre-tokenization;specialonly controlsskip_special_tokenson decode. ModernBERT has 109 such tokens (whitespace runs,|||IP_ADDRESS|||):" leading"→ reference[50275, 16378], engine[245, 4283, …].vocabSizealso reports 50280 instead of 50368. - The
normalizerblock is ignored. ModernBERT declares{"type": "NFC"}; a decomposede+ U+0301 is encoded as three byte tokens instead of the composedétoken (reference[…, 3974], engine[…, 299, 136, 212]). OnlySentencePieceTokenizerlooks at the block, and only to sniffPrepend(SentencePieceTokenizer.kt#L382-L391). - Metaspace BPE with byte fallback.
model.type: "BPE"+pre_tokenizer: Metaspace+byte_fallback: trueis the Llama/Gemma/mmBERT family. It is routed to the byte-level class, which applies the GPT-2 regex and the byte→unicode table to a▁vocabulary. This is the encode-side twin of #912 (which reports the decode-side crash on the same files).
Who benefits: anyone loading a SafeTensors checkpoint with its tokenizer.json — encoder models (ModernBERT, mmBERT, BERT-class rerankers) and every decoder exported with a current tokenizers. The downstream SKaiNET-transformers factory delegates model.type: BPE straight here (UpstreamTokenizerAdapter), so the fix lands in one place.
Summary (1–2 sentences):
Make
TokenizerFactory.fromTokenizerJsonreproduce the Hugging Facetokenizersids for BPEtokenizer.jsonfiles: accept array-form merges, match all added tokens, apply the declared normalizer, and dispatch Metaspace BPE to a SentencePiece-style BPE encoder instead of the byte-level one.
🔍 A: ASSESS — Feasibility & Impact
✔️ Feasibility
- Gaps 1 and 2 are local, few-line changes with existing synthetic tests to extend (
TokenizerFactoryDispatchTest). Both are good first issues. - Gap 3 needs an NFC/NFKC implementation reachable from
commonMain(JVM hasjava.text.Normalizer; Kotlin/Native, JS and Wasm do not ship one). Whether to goexpect/actual, table-driven, or JVM-first-with-passthrough is the one design decision in this feature → research lane first. - Gap 4 is the largest: the merge-rank BPE core in
QwenByteLevelBpeTokenizeris correct, but its alphabet mapping (bytes → unicode, GPT-2 regex) is hard-wired. Splitting the BPE core from the pre-tokenizer/alphabet step lets one core serve ByteLevel and Metaspace.SentencePieceTokenizeris not a substitute: it implements llama.cpp score-priority merging over a Unigram vocabulary and expects[token, score]pairs.
✔️ Expected Impact
- Every
tokenizer.jsonexported bytokenizers≥ 0.20 loads again (gap 1 alone unblocks this). - Encoder-class checkpoints become loadable without a per-model tokenizer fork.
- Closes the encode half of #912; the decode half (
decodethrowing on▁, bos/eos fromadded_tokens) is finished by the same lane.
✔️ Risks / Constraints
- Changing
vocabSizeto include added tokens changes a public value; callers that size embedding tables from it get the correct number afterwards, but it is a behaviour change — list it in the changelog. - Added-token matching must respect
lstrip/rstrip/single_word/normalizedflags; the Laya files use onlynormalized: trueon the non-special entries. Anything else → separate issue, do not block. - NFC on non-JVM targets: if the research lane finds no cheap table, ship JVM
actual+ a documented passthrough elsewhere rather than blocking the whole feature.
✔️ Dependencies
skainet-io-coreonly (sk.ainet.io.tokenizer). No new third-party library unless the NFC research recommends one (it should not).- Reference oracle: Hugging Face
tokenizers(Python), pinned per golden file.
📚 R: RESEARCH — What Must Be Understood First?
Research Tasks
- Survey which
normalizer/pre_tokenizer/decodercombinations appear in thetokenizer.jsonof popular public checkpoints (Lane 1). - Decide the NFC strategy for
commonMain(Lane 1 output, consumed by the normalizer lane). - Pin the exact matching semantics of
added_tokensflags intokenizers(AddedVocabulary), with citations.
Open Questions
- Should
vocabSizeinclude added tokens whose id is abovemodel.vocab? (Reference: yes —get_vocab_size(with_added_tokens=True).)- Does the
ignore_merges: trueflag (Llama 3 files) need handling in the same pass as gap 4?
🛠️ C: CONTRIBUTE — Implementation Plan
Development Tasks
- Lane 1 · numerics — normalizer / pre-tokenizer survey and NFC recommendation
- Lane 2 · kotlin-core — array-form
model.merges(good first issue) - Lane 2 · kotlin-core — atomic non-special
added_tokens+vocabSize(good first issue) - Lane 2 · kotlin-core — apply the
normalizerblock (blocked by Lane 1) - Lane 2 · kotlin-core — Metaspace BPE with byte fallback (encode side of #912)
- Lane 4 · numerics — golden encode fixtures from
tokenizersfor the two Laya files (good first issue) - Lane 4 · kotlin-core — fixture-gated jvmTest replaying the goldens (good first issue, blocked by the golden lane)
- Lane 5 · docs —
tokenizer.jsonsupport matrix (good first issue) - Lane 6 · review — DARC review
Acceptance Criteria
- Both Laya
tokenizer.jsonfiles load as published (no merge rewriting) and encode the fixture texts byte-identically to the currenttokenizersrelease (0.23.3 at filing time) (add_special_tokens=False). -
vocabSizeequalsget_vocab_size(with_added_tokens=True). - The #912 decode fixture passes on JVM and on
linuxX64Test. - Existing Qwen fixture tests unchanged.
- No regressions elsewhere; reviewed by someone who did not implement the lanes.
💬 Additional Notes
Fixture texts (shared by every lane; reference ids are produced with tokenizers 0.23.3 (identical under 0.20.3), Tokenizer.from_file(path).encode(text, add_special_tokens=False).ids):
Bitte die doppelte Rechnung prüfen.
Which department handles this request?
Straße, Übergrößen und Ärger – naïve café façade 😀👍🏽 é ← the final "é" is e + U+0301 (decomposed)
false: No duplicate is mentioned
leading spaces and 12345 numbers, hyphen-ated words don't
Reference files (Apache-2.0, public): ModernBERT tokenizer/tokenizer.json — normalizer: NFC, pre_tokenizer: ByteLevel, 50,280 vocab + 116 added tokens (7 special); mmBERT multilingual/tokenizer/tokenizer.json — normalizer: Replace(" "→"▁"), pre_tokenizer: Metaspace(prepend_scheme: always), byte_fallback: true, 256,000 vocab, 580,604 merges, 255 <0xNN> tokens.
Lane breakdown
| Lane | Applies? | Why |
|---|---|---|
| 0 Design (SKEEP) | no | Additive behaviour behind the existing Tokenizer interface and factory; no API shape change. If the Metaspace lane ends up needing a new public class, flag it there. |
| 1 Numerics / Research | yes | Normalizer survey + NFC strategy; Python only. |
| 2 Kotlin core | yes | Four tasks, two of them first issues. |
| 3 Platform verification | folded in | allTests and linuxX64Test already run skainet-io-core in CI; the fixture lane asserts on both. |
| 4 Ground truth / CI | yes | Python goldens + Kotlin replay test. |
| 5 Docs | yes | Support matrix page; the engine docs have no tokenizer page today. |
| 6 Review | yes |
Related: #912 (decode side, native tests), #858 (legacy files without model.type), #463 (original byte-level BPE fix). Downstream consumer: SKaiNET-transformers TokenizerFactory.fromTokenizerJsonString, which delegates model.type: BPE here.
- Lenguaje dominante
- Kotlin
- Estrellas
- 52
- Forks
- 15
- Merge medio
- 1 d 15 h
- PR fusionados (30 d)
- 36
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de SKaiNET-developers/SKaiNET
-
coding good first issue size:xs skill:kotlin-core sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
SKaiNET-developers/SKaiNET#1323 ·
Los mantenedores suelen responder en 1 día
-
coding good first issue platform size:xs skill:js sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
SKaiNET-developers/SKaiNET#1232 ·
Los mantenedores suelen responder en 1 día
-
skeep tracking
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
SKaiNET-developers/SKaiNET#1331 ·
Los mantenedores suelen responder en 1 día
-
assessment size:s skill:review sub-issue
Dificultad 4/5 1-2 días Aptitud para principiantes 18/100
SKaiNET-developers/SKaiNET#1330 ·
Los mantenedores suelen responder en 1 día
-
documentation good first issue size:s skill:docs sub-issue
Dificultad 3/5 1-2 días Aptitud para principiantes 78/100
SKaiNET-developers/SKaiNET#1329 ·
Los mantenedores suelen responder en 1 día
Todos los issues de SKaiNET-developers/SKaiNET
Issues similares
-
[Submission] 抖音火山版Abiertosubmit-adaption submit-adaption-pre
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
BetterAndroid/android-notification-icon-project#744 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 64/100
utopia-rise/godot-jvm#1004 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
pedroSG94/RootEncoder#2213 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
[Feature]: AC charger voltageAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
dzid26/TeslaBatteryBLE#183 ·
Los mantenedores suelen responder en 1 día