Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Lane 1 · numerics] tokenizer.json parity — normalizer / pre-tokenizer survey and NFC strategy for commonMain

Đang mở
#1,322 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
57/100
Loại issue
Tính năng
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
huggingface, kotlin, markdown, python
Lĩnh vực
data, machine-learning

Hướng nghiên cứu

Start by reading the two linked ModernBERT and mmBERT tokenizer.json files and the Hugging Face tokenizers documentation, then use a small Python script with json.load to survey 15–20 checkpoint files. Record the requested tokenizer fields in a Markdown table with exact revision links, summarize normalizer and pre-tokenizer types and merge formats, and compare NFC options using the linked Unicode data. Done means posting the completed table and recommendation as a comment on #1321.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

good first issue research size:s skill:numerics sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 1 · Numerics / Research
Skill needed: Python and the Hugging Face tokenizers library. No Kotlin, no SKaiNET codebase knowledge.
Size: s (an afternoon)
Blocked by: nothing

What to do

  1. Pick 15–20 public checkpoints that ship a tokenizer.json and cover the families SKaiNET loads or plans to load: GPT-2, Qwen2.5/Qwen3, Llama 3.x, Gemma 3/4, Mistral (HF export), SmolLM2, BitNet, ModernBERT, mmBERT, BERT-base, XLM-R, T5, Whisper, Moonshine.
  2. For each file record: model.type, model.byte_fallback, model.ignore_merges, the form of model.merges[0] (string or array), the full normalizer block, the pre_tokenizer block, the decoder block, and the number of added_tokens with special: false. A 30-line script with json.load is enough — no need to run the tokenizers.
  3. Produce one Markdown table, sorted by family, and a short section answering: which normalizer types occur at all (expected: NFC, NFKC, Replace, Prepend, Sequence, BertNormalizer, Lowercase, StripAccents), which pre-tokenizer types (ByteLevel, Metaspace, Split, Sequence, BertPreTokenizer), and how many files use the array form of merges.
  4. Recommend an NFC strategy for Kotlin commonMain: (a) expect/actual with java.text.Normalizer on JVM and String.prototype.normalize on JS/Wasm, with a passthrough on Kotlin/Native; (b) a table-driven implementation (quote the table size for the composition exclusions + canonical compositions); or (c) something better. One paragraph per option, one recommendation.
  5. Post the table and recommendation as a comment on #1321.

Acceptance

  • Table with ≥ 15 checkpoints, each row linking the exact tokenizer.json revision on the Hub
  • Count of files whose merges are in array form vs string form
  • NFC recommendation with a link to the Unicode data it would depend on
  • Result reported back on #1321

Notes

The two files that triggered the parent are ModernBERT and mmBERT; include both so the Kotlin lanes can cross-check. If you find a normalizer type that cannot be implemented without a large Unicode table (e.g. NFKC), say so — that is a finding, not a blocker.

Ngôn ngữ chính
Kotlin
Star
52
Fork
15
Merge trung bình
1 ngày 15 giờ
Pull request đã merge (30 ngày)
36

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của SKaiNET-developers/SKaiNET

Tất cả issue của SKaiNET-developers/SKaiNET

Issue tương tự

Thêm issue về Kotlin

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.