Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Feature]: HF tokenizer.json parity for BPE encoders — array-form merges, normalizer, non-special added tokens, Metaspace BPE

Abierto
#1,321 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
kotlin, python
Área
ai, backend

Línea de trabajo

Start by reading TokenizerFactory.kt, QwenByteLevelBpeTokenizer.kt, and the existing TokenizerFactoryDispatchTest; compare the documented merge parsing and added token handling with Hugging Face tokenizers 0.23.3. The issue divides the work into separate lanes, with NFC strategy and Metaspace BPE requiring research and design. Done means both published Laya tokenizer files load without rewriting and match the reference fixtures, with the stated vocabulary size and decode checks.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

darc enhancement tracking

🧠 D: DEFINE — Problem & Opportunity

sk.ainet.io.tokenizer.TokenizerFactory.fromTokenizerJson (skainet-io-core 0.57.0) cannot load the tokenizer.json of any current Hugging Face encoder checkpoint, and when the file is patched so that it loads, the ids it produces differ from the tokenizers reference. Measured on the two tokenizer files shipped with the Laya checkpoints — ModernBERT-large (byte-level BPE, NFC normalizer) and mmBERT-base (SentencePiece-style BPE, Metaspace pre-tokenizer, byte fallback):

tokenizer.json As published With model.merges rewritten to the legacy "a b" string form
ModernBERT (tokenizer/tokenizer.json) load fails: IllegalArgumentException: Element class kotlinx.serialization.json.JsonArray … is not a JsonPrimitive QwenByteLevelBpeTokenizer, 3 of 5 sample texts identical to the reference
mmBERT (multilingual/tokenizer/tokenizer.json) same load failure QwenByteLevelBpeTokenizer, 0 of 5 identical (byte-level pre-tokenization applied to a ▁ vocabulary)

Four independent gaps, in increasing size:

  1. Array form of model.merges. the Hugging Face tokenizers Python library (PyPI; not a SKaiNET version) writes merges since its 0.20 release as [["a","b"], …] (verified with 0.20.3 and with the current 0.23.3: both write the array form and still read the string form). The builder at QwenByteLevelBpeTokenizer.kt#L243-L255 only accepts "a b" strings, so every freshly exported tokenizer fails to load.
  2. Non-special added_tokens are dropped. The special: false entries are filtered out at QwenByteLevelBpeTokenizer.kt#L257-L267 and TokenizerFactory.kt#L155-L169. The reference matches every added token atomically before pre-tokenization; special only controls skip_special_tokens on decode. ModernBERT has 109 such tokens (whitespace runs, |||IP_ADDRESS|||): " leading" → reference [50275, 16378], engine [245, 4283, …]. vocabSize also reports 50280 instead of 50368.
  3. The normalizer block is ignored. ModernBERT declares {"type": "NFC"}; a decomposed e + U+0301 is encoded as three byte tokens instead of the composed é token (reference […, 3974], engine […, 299, 136, 212]). Only SentencePieceTokenizer looks at the block, and only to sniff Prepend (SentencePieceTokenizer.kt#L382-L391).
  4. Metaspace BPE with byte fallback. model.type: "BPE" + pre_tokenizer: Metaspace + byte_fallback: true is the Llama/Gemma/mmBERT family. It is routed to the byte-level class, which applies the GPT-2 regex and the byte→unicode table to a ▁ vocabulary. This is the encode-side twin of #912 (which reports the decode-side crash on the same files).

Who benefits: anyone loading a SafeTensors checkpoint with its tokenizer.json — encoder models (ModernBERT, mmBERT, BERT-class rerankers) and every decoder exported with a current tokenizers. The downstream SKaiNET-transformers factory delegates model.type: BPE straight here (UpstreamTokenizerAdapter), so the fix lands in one place.

Summary (1–2 sentences):

Make TokenizerFactory.fromTokenizerJson reproduce the Hugging Face tokenizers ids for BPE tokenizer.json files: accept array-form merges, match all added tokens, apply the declared normalizer, and dispatch Metaspace BPE to a SentencePiece-style BPE encoder instead of the byte-level one.


🔍 A: ASSESS — Feasibility & Impact

✔️ Feasibility
  • Gaps 1 and 2 are local, few-line changes with existing synthetic tests to extend (TokenizerFactoryDispatchTest). Both are good first issues.
  • Gap 3 needs an NFC/NFKC implementation reachable from commonMain (JVM has java.text.Normalizer; Kotlin/Native, JS and Wasm do not ship one). Whether to go expect/actual, table-driven, or JVM-first-with-passthrough is the one design decision in this feature → research lane first.
  • Gap 4 is the largest: the merge-rank BPE core in QwenByteLevelBpeTokenizer is correct, but its alphabet mapping (bytes → unicode, GPT-2 regex) is hard-wired. Splitting the BPE core from the pre-tokenizer/alphabet step lets one core serve ByteLevel and Metaspace. SentencePieceTokenizer is not a substitute: it implements llama.cpp score-priority merging over a Unigram vocabulary and expects [token, score] pairs.
✔️ Expected Impact
  • Every tokenizer.json exported by tokenizers ≥ 0.20 loads again (gap 1 alone unblocks this).
  • Encoder-class checkpoints become loadable without a per-model tokenizer fork.
  • Closes the encode half of #912; the decode half (decode throwing on ▁, bos/eos from added_tokens) is finished by the same lane.
✔️ Risks / Constraints
  • Changing vocabSize to include added tokens changes a public value; callers that size embedding tables from it get the correct number afterwards, but it is a behaviour change — list it in the changelog.
  • Added-token matching must respect lstrip/rstrip/single_word/normalized flags; the Laya files use only normalized: true on the non-special entries. Anything else → separate issue, do not block.
  • NFC on non-JVM targets: if the research lane finds no cheap table, ship JVM actual + a documented passthrough elsewhere rather than blocking the whole feature.
✔️ Dependencies
  • skainet-io-core only (sk.ainet.io.tokenizer). No new third-party library unless the NFC research recommends one (it should not).
  • Reference oracle: Hugging Face tokenizers (Python), pinned per golden file.

📚 R: RESEARCH — What Must Be Understood First?

Research Tasks
  • Survey which normalizer / pre_tokenizer / decoder combinations appear in the tokenizer.json of popular public checkpoints (Lane 1).
  • Decide the NFC strategy for commonMain (Lane 1 output, consumed by the normalizer lane).
  • Pin the exact matching semantics of added_tokens flags in tokenizers (AddedVocabulary), with citations.
Open Questions
  • Should vocabSize include added tokens whose id is above model.vocab? (Reference: yes — get_vocab_size(with_added_tokens=True).)
  • Does the ignore_merges: true flag (Llama 3 files) need handling in the same pass as gap 4?

🛠️ C: CONTRIBUTE — Implementation Plan

Development Tasks
  • Lane 1 · numerics — normalizer / pre-tokenizer survey and NFC recommendation
  • Lane 2 · kotlin-core — array-form model.merges (good first issue)
  • Lane 2 · kotlin-core — atomic non-special added_tokens + vocabSize (good first issue)
  • Lane 2 · kotlin-core — apply the normalizer block (blocked by Lane 1)
  • Lane 2 · kotlin-core — Metaspace BPE with byte fallback (encode side of #912)
  • Lane 4 · numerics — golden encode fixtures from tokenizers for the two Laya files (good first issue)
  • Lane 4 · kotlin-core — fixture-gated jvmTest replaying the goldens (good first issue, blocked by the golden lane)
  • Lane 5 · docs — tokenizer.json support matrix (good first issue)
  • Lane 6 · review — DARC review
Acceptance Criteria
  • Both Laya tokenizer.json files load as published (no merge rewriting) and encode the fixture texts byte-identically to the current tokenizers release (0.23.3 at filing time) (add_special_tokens=False).
  • vocabSize equals get_vocab_size(with_added_tokens=True).
  • The #912 decode fixture passes on JVM and on linuxX64Test.
  • Existing Qwen fixture tests unchanged.
  • No regressions elsewhere; reviewed by someone who did not implement the lanes.

💬 Additional Notes

Fixture texts (shared by every lane; reference ids are produced with tokenizers 0.23.3 (identical under 0.20.3), Tokenizer.from_file(path).encode(text, add_special_tokens=False).ids):

Bitte die doppelte Rechnung prüfen.
Which department handles this request?
Straße, Übergrößen und Ärger – naïve café façade 😀👍🏽 é      ← the final "é" is e + U+0301 (decomposed)
false: No duplicate is mentioned
   leading spaces and 12345 numbers, hyphen-ated words don't

Reference files (Apache-2.0, public): ModernBERT tokenizer/tokenizer.json — normalizer: NFC, pre_tokenizer: ByteLevel, 50,280 vocab + 116 added tokens (7 special); mmBERT multilingual/tokenizer/tokenizer.json — normalizer: Replace(" "→"▁"), pre_tokenizer: Metaspace(prepend_scheme: always), byte_fallback: true, 256,000 vocab, 580,604 merges, 255 <0xNN> tokens.

Lane breakdown

Lane Applies? Why
0 Design (SKEEP) no Additive behaviour behind the existing Tokenizer interface and factory; no API shape change. If the Metaspace lane ends up needing a new public class, flag it there.
1 Numerics / Research yes Normalizer survey + NFC strategy; Python only.
2 Kotlin core yes Four tasks, two of them first issues.
3 Platform verification folded in allTests and linuxX64Test already run skainet-io-core in CI; the fixture lane asserts on both.
4 Ground truth / CI yes Python goldens + Kotlin replay test.
5 Docs yes Support matrix page; the engine docs have no tokenizer page today.
6 Review yes

Related: #912 (decode side, native tests), #858 (legacy files without model.type), #463 (original byte-level BPE fix). Downstream consumer: SKaiNET-transformers TokenizerFactory.fromTokenizerJsonString, which delegates model.type: BPE here.

Lenguaje dominante
Kotlin
Estrellas
52
Forks
15
Merge medio
1 d 15 h
PR fusionados (30 d)
36

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de SKaiNET-developers/SKaiNET

Todos los issues de SKaiNET-developers/SKaiNET

Issues similares

Más issues de Kotlin

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.