Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Lane 1 · numerics] tokenizer.json parity — normalizer / pre-tokenizer survey and NFC strategy for commonMain

オープン
#1,322 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
1〜2日
初心者へのやさしさ
57/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
huggingface, kotlin, markdown, python

調査の方向性

Start by reading the two linked ModernBERT and mmBERT tokenizer.json files and the Hugging Face tokenizers documentation, then use a small Python script with json.load to survey 15–20 checkpoint files. Record the requested tokenizer fields in a Markdown table with exact revision links, summarize normalizer and pre-tokenizer types and merge formats, and compare NFC options using the linked Unicode data. Done means posting the completed table and recommendation as a comment on #1321.

索引モデルが issue の本文から書いたものです。

説明

good first issue research size:s skill:numerics sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 1 · Numerics / Research
Skill needed: Python and the Hugging Face tokenizers library. No Kotlin, no SKaiNET codebase knowledge.
Size: s (an afternoon)
Blocked by: nothing

What to do

  1. Pick 15–20 public checkpoints that ship a tokenizer.json and cover the families SKaiNET loads or plans to load: GPT-2, Qwen2.5/Qwen3, Llama 3.x, Gemma 3/4, Mistral (HF export), SmolLM2, BitNet, ModernBERT, mmBERT, BERT-base, XLM-R, T5, Whisper, Moonshine.
  2. For each file record: model.type, model.byte_fallback, model.ignore_merges, the form of model.merges[0] (string or array), the full normalizer block, the pre_tokenizer block, the decoder block, and the number of added_tokens with special: false. A 30-line script with json.load is enough — no need to run the tokenizers.
  3. Produce one Markdown table, sorted by family, and a short section answering: which normalizer types occur at all (expected: NFC, NFKC, Replace, Prepend, Sequence, BertNormalizer, Lowercase, StripAccents), which pre-tokenizer types (ByteLevel, Metaspace, Split, Sequence, BertPreTokenizer), and how many files use the array form of merges.
  4. Recommend an NFC strategy for Kotlin commonMain: (a) expect/actual with java.text.Normalizer on JVM and String.prototype.normalize on JS/Wasm, with a passthrough on Kotlin/Native; (b) a table-driven implementation (quote the table size for the composition exclusions + canonical compositions); or (c) something better. One paragraph per option, one recommendation.
  5. Post the table and recommendation as a comment on #1321.

Acceptance

  • Table with ≥ 15 checkpoints, each row linking the exact tokenizer.json revision on the Hub
  • Count of files whose merges are in array form vs string form
  • NFC recommendation with a link to the Unicode data it would depend on
  • Result reported back on #1321

Notes

The two files that triggered the parent are ModernBERT and mmBERT; include both so the Kotlin lanes can cross-check. If you find a normalizer type that cannot be implemented without a large Unicode table (e.g. NFKC), say so — that is a finding, not a blocker.

主要言語
Kotlin
スター
52
フォーク
15
平均マージ
1日 15時間
マージ済み PR(30日)
36

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

SKaiNET-developers/SKaiNET のほかの issue

SKaiNET-developers/SKaiNET の issue をすべて見る

似ている issue

Kotlin の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。