Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Lane 4 · numerics] tokenizer.json parity — golden encode fixtures from Hugging Face tokenizers for the ModernBERT and mmBERT files

オープン
#1,327 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
76/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python

調査の方向性

Create scripts/tokenizer_goldens.py; use the five reference texts from #1321 and the listed edge cases, then generate goldens for both Laya tokenizer paths at the specified revision. Start by checking #1321 for its five expected rows and the Qwen fixture test header for the command and pip pin format. Done means deterministic JSON outputs with source SHA-256 and tokenizers version committed under the stated resources directory.

索引モデルが issue の本文から書いたものです。

説明

coding good first issue size:xs skill:numerics sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 4 · Ground truth / CI
Skill needed: Python (tokenizers, huggingface_hub). No Kotlin, no SKaiNET knowledge.
Size: xs–s (an hour or two)
Blocked by: nothing

What to do

  1. Write scripts/tokenizer_goldens.py (new file) that, given a Hub repo id, a revision and a path to a tokenizer.json, downloads the file, computes its SHA-256, and encodes a fixed list of texts with Tokenizer.from_file(...).encode(text, add_special_tokens=False). Output one JSON per tokenizer:
    {"repo": "convaiinnovations/laya", "revision": "7b928d82…", "path": "tokenizer/tokenizer.json",
     "sha256": "…", "tokenizers_version": "0.23.3", "vocab_size_with_added": 50368,
     "cases": [{"text": "Bitte die doppelte Rechnung prüfen.", "ids": [12871, 442, …], "tokens": ["Bit", "te", …]}]}
    
  2. Texts: the five fixture texts from #1321 plus ~15 more that each target one feature — a lone special token (<|endoftext|>, <bos>), a non-special added token, tabs and newlines, a run of 24 spaces (ModernBERT has an added token for it), CJK, Arabic, an emoji with skin-tone modifier, decomposed vs composed accents, digits (1234567890), the empty string, a single space.
  3. Generate goldens for both Laya files: tokenizer/tokenizer.json and multilingual/tokenizer/tokenizer.json at revision 7b928d828b7b0e022f929d9bd2e44165aa270148 of https://huggingface.co/convaiinnovations/laya (Apache-2.0).
  4. Commit the two JSON files under skainet-io/skainet-io-core/src/jvmTest/resources/tokenizer-goldens/ (they are a few KB; the tokenizer files themselves are not committed — the Kotlin lane downloads them like the existing Qwen fixtures).
  5. Put the exact command line and the pip pin in the script header, in the style of the header comment of QwenByteLevelBpeTokenizerFixtureTest.kt.

Acceptance

  • Script reproduces the five reference rows quoted in #1321 byte-for-byte
  • Two golden JSON files committed, each with sha256 of the source tokenizer.json and the tokenizers version
  • Running the script twice gives identical output (deterministic ordering)
  • Result reported back on #1321

Notes

Keep the script free of SKaiNET-specific logic so the sibling repository (SKaiNET-transformers) can reuse it for its own model families. If a text makes tokenizers itself warn or behave oddly, keep the text and note it in the JSON ("note" field) — surprising oracle behaviour is exactly what the Kotlin side must match.

主要言語
Kotlin
スター
52
フォーク
15
平均マージ
1日 15時間
マージ済み PR(30日)
36

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

SKaiNET-developers/SKaiNET のほかの issue

SKaiNET-developers/SKaiNET の issue をすべて見る

似ている issue

Kotlin の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。