Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Lane 4 · numerics] tokenizer.json parity — golden encode fixtures from Hugging Face tokenizers for the ModernBERT and mmBERT files

Open
#1,327 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
76/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Create scripts/tokenizer_goldens.py; use the five reference texts from #1321 and the listed edge cases, then generate goldens for both Laya tokenizer paths at the specified revision. Start by checking #1321 for its five expected rows and the Qwen fixture test header for the command and pip pin format. Done means deterministic JSON outputs with source SHA-256 and tokenizers version committed under the stated resources directory.

Written by the indexing model from the issue text.

Description

coding good first issue size:xs skill:numerics sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 4 · Ground truth / CI
Skill needed: Python (tokenizers, huggingface_hub). No Kotlin, no SKaiNET knowledge.
Size: xs–s (an hour or two)
Blocked by: nothing

What to do

  1. Write scripts/tokenizer_goldens.py (new file) that, given a Hub repo id, a revision and a path to a tokenizer.json, downloads the file, computes its SHA-256, and encodes a fixed list of texts with Tokenizer.from_file(...).encode(text, add_special_tokens=False). Output one JSON per tokenizer:
    {"repo": "convaiinnovations/laya", "revision": "7b928d82…", "path": "tokenizer/tokenizer.json",
     "sha256": "…", "tokenizers_version": "0.23.3", "vocab_size_with_added": 50368,
     "cases": [{"text": "Bitte die doppelte Rechnung prüfen.", "ids": [12871, 442, …], "tokens": ["Bit", "te", …]}]}
    
  2. Texts: the five fixture texts from #1321 plus ~15 more that each target one feature — a lone special token (<|endoftext|>, <bos>), a non-special added token, tabs and newlines, a run of 24 spaces (ModernBERT has an added token for it), CJK, Arabic, an emoji with skin-tone modifier, decomposed vs composed accents, digits (1234567890), the empty string, a single space.
  3. Generate goldens for both Laya files: tokenizer/tokenizer.json and multilingual/tokenizer/tokenizer.json at revision 7b928d828b7b0e022f929d9bd2e44165aa270148 of https://huggingface.co/convaiinnovations/laya (Apache-2.0).
  4. Commit the two JSON files under skainet-io/skainet-io-core/src/jvmTest/resources/tokenizer-goldens/ (they are a few KB; the tokenizer files themselves are not committed — the Kotlin lane downloads them like the existing Qwen fixtures).
  5. Put the exact command line and the pip pin in the script header, in the style of the header comment of QwenByteLevelBpeTokenizerFixtureTest.kt.

Acceptance

  • Script reproduces the five reference rows quoted in #1321 byte-for-byte
  • Two golden JSON files committed, each with sha256 of the source tokenizer.json and the tokenizers version
  • Running the script twice gives identical output (deterministic ordering)
  • Result reported back on #1321

Notes

Keep the script free of SKaiNET-specific logic so the sibling repository (SKaiNET-transformers) can reuse it for its own model families. If a text makes tokenizers itself warn or behave oddly, keep the text and note it in the JSON ("note" field) — surprising oracle behaviour is exactly what the Kotlin side must match.

Dominant language
Kotlin
Stars
52
Forks
15
Avg merge
1d 15h
Merged PRs (30d)
36

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SKaiNET-developers/SKaiNET

All issues in SKaiNET-developers/SKaiNET

Similar issues

More Kotlin issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.