Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Lane 5 · docs] tokenizer.json parity — reference page: tokenizer.json support matrix

Open
#1,329 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
78/100
Issue type
Documentation
Clarity
Clearly specified
Activity status
Active
Tech stack
kotlin
Domain
ai, documentation

Research direction

Read TokenizerFactory.kt and the Lane 2 measurements from #1321 to fill the support matrix, then follow an existing reference page for its Antora structure. Add tokenizer-json-support.adoc, register it in nav.adoc, and add the KDoc pointer; check the Lane 4 sub-issues for the Python golden script and gated test. Done when the page is listed under Reference, unsupported or partial rows link to open issues, and the result is reported on #1321.

Written by the indexing model from the issue text.

Description

documentation good first issue size:s skill:docs sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 5 · Docs
Skill needed: AsciiDoc and the ability to read a Kotlin when statement. No build or tokenizer knowledge.
Size: s (a few hours)
Blocked by: nothing for the first version; update the table as the Lane 2 PRs merge.

What to do

  1. The engine docs have no tokenizer page. Add docs/modules/ROOT/pages/reference/tokenizer-json-support.adoc and register it in docs/modules/ROOT/nav.adoc under Reference.
  2. Content: a support matrix of what sk.ainet.io.tokenizer.TokenizerFactory accepts, one row per tokenizer.json feature — model.type (BPE / Unigram / WordPiece), model.merges form (string / array), byte_fallback, ignore_merges, each normalizer type, each pre_tokenizer type, added_tokens (special: true / false, the four flags), post_processor (not applied — callers add BOS/EOS themselves, say so), decoder. Columns: Supported (yes / no / partial), Since version, Issue link. Derive the current state from TokenizerFactory.kt and the #1321 measurements; link each "no" to its lane sub-issue.
  3. A second short section "GGUF vs tokenizer.json": the factory dispatches per architecture, not per file format (the KDoc at the top of TokenizerFactory.kt already says this well — reuse it).
  4. A "How to verify parity" paragraph pointing to the Python golden script and the gated test from the Lane 4 sub-issues.
  5. Add a one-line pointer from the TokenizerFactory KDoc to the new page. Build the docs locally if you can (docs/README or the Antora playbook in the repo explains how); otherwise a careful preview in the GitHub editor is acceptable for a first PR.

Acceptance

  • Page renders in the Antora nav under Reference
  • Every "no" / "partial" row links to an open issue
  • fromTokenizerJson KDoc links to the page
  • Result reported back on #1321

Notes

Follow the structure of an existing reference page in docs/modules/ROOT/pages/reference/ for headings and the :description: attribute. The downstream repo (SKaiNET-transformers) has explanation/tokenizer-internals.adoc; do not duplicate it, link to it for the algorithm descriptions.

Dominant language
Kotlin
Stars
52
Forks
15
Avg merge
1d 15h
Merged PRs (30d)
36

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SKaiNET-developers/SKaiNET

All issues in SKaiNET-developers/SKaiNET

Similar issues

More Kotlin issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.