Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Lane 5 · docs] tokenizer.json parity — reference page: tokenizer.json support matrix

オープン
#1,329 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
78/100
issue の種類
ドキュメント
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
kotlin
領域
ai, documentation

調査の方向性

Read TokenizerFactory.kt and the Lane 2 measurements from #1321 to fill the support matrix, then follow an existing reference page for its Antora structure. Add tokenizer-json-support.adoc, register it in nav.adoc, and add the KDoc pointer; check the Lane 4 sub-issues for the Python golden script and gated test. Done when the page is listed under Reference, unsupported or partial rows link to open issues, and the result is reported on #1321.

索引モデルが issue の本文から書いたものです。

説明

documentation good first issue size:s skill:docs sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 5 · Docs
Skill needed: AsciiDoc and the ability to read a Kotlin when statement. No build or tokenizer knowledge.
Size: s (a few hours)
Blocked by: nothing for the first version; update the table as the Lane 2 PRs merge.

What to do

  1. The engine docs have no tokenizer page. Add docs/modules/ROOT/pages/reference/tokenizer-json-support.adoc and register it in docs/modules/ROOT/nav.adoc under Reference.
  2. Content: a support matrix of what sk.ainet.io.tokenizer.TokenizerFactory accepts, one row per tokenizer.json feature — model.type (BPE / Unigram / WordPiece), model.merges form (string / array), byte_fallback, ignore_merges, each normalizer type, each pre_tokenizer type, added_tokens (special: true / false, the four flags), post_processor (not applied — callers add BOS/EOS themselves, say so), decoder. Columns: Supported (yes / no / partial), Since version, Issue link. Derive the current state from TokenizerFactory.kt and the #1321 measurements; link each "no" to its lane sub-issue.
  3. A second short section "GGUF vs tokenizer.json": the factory dispatches per architecture, not per file format (the KDoc at the top of TokenizerFactory.kt already says this well — reuse it).
  4. A "How to verify parity" paragraph pointing to the Python golden script and the gated test from the Lane 4 sub-issues.
  5. Add a one-line pointer from the TokenizerFactory KDoc to the new page. Build the docs locally if you can (docs/README or the Antora playbook in the repo explains how); otherwise a careful preview in the GitHub editor is acceptable for a first PR.

Acceptance

  • Page renders in the Antora nav under Reference
  • Every "no" / "partial" row links to an open issue
  • fromTokenizerJson KDoc links to the page
  • Result reported back on #1321

Notes

Follow the structure of an existing reference page in docs/modules/ROOT/pages/reference/ for headings and the :description: attribute. The downstream repo (SKaiNET-transformers) has explanation/tokenizer-internals.adoc; do not duplicate it, link to it for the algorithm descriptions.

主要言語
Kotlin
スター
52
フォーク
15
平均マージ
1日 15時間
マージ済み PR(30日)
36

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

SKaiNET-developers/SKaiNET のほかの issue

SKaiNET-developers/SKaiNET の issue をすべて見る

似ている issue

Kotlin の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。