Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Lane 2 · kotlin-core] tokenizer.json parity — match non-special added_tokens atomically and count them in vocabSize

Abierto
#1,324 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-3 horas
Aptitud para principiantes
72/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
kotlin
Área
tooling

Línea de trabajo

Read the added-token loop in QwenByteLevelBpeTokenizer.kt and TokenizerFactory.kt, then inspect the related tests in TokenizerFactoryDispatchTest.kt. Add a synthetic case where a non-special token would otherwise split, and check that it encodes atomically and vocabSize includes its id. Run ./gradlew :skainet-io:skainet-io-core:jvmTest; done means the test passes and the changelog records the vocabSize change.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

coding good first issue size:s skill:kotlin-core sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 2 · Kotlin core
Skill needed: Kotlin. Helpful to have read how tokenizers matches added tokens (one page of docs, linked below). No tensor knowledge.
Size: s (a few hours)
Blocked by: nothing (independent of the merges lane; both touch the same file, so rebase once)

What to do

  1. Read how the reference matches added tokens: every entry of added_tokens is matched atomically, before pre-tokenization, regardless of special. The special flag only decides whether the token is dropped by decode(skip_special_tokens=True). See the AddedVocabulary section of the tokenizers docs: https://huggingface.co/docs/tokenizers/api/added-tokens
  2. In QwenByteLevelBpeTokenizer.kt#L257-L267 the loop skips entries with "special": false. Register all entries in the atomic-match map. Keep a separate Set<Int> of the ids that are marked special so a later change can implement skip_special_tokens; nothing else needs that set today.
  3. Make the same change in TokenizerFactory.kt#L155-L169 for the Unigram path (wrapSentencePieceWithSpecialsFromJson).
  4. Added tokens may have ids above model.vocab (ModernBERT: [UNK]=50280 … and the whitespace runs up to 50367). Extend the tokens array so vocabSize equals the reference get_vocab_size(with_added_tokens=True) — 50368 for ModernBERT instead of today's 50280 — and so decode can return their text.
  5. Tests, in TokenizerFactoryDispatchTest.kt: a synthetic file with one special: false added token whose content would otherwise BPE-split; assert it encodes to its id and that vocabSize counts it. Run ./gradlew :skainet-io:skainet-io-core:jvmTest.

Acceptance

  • Synthetic test: non-special added token encodes atomically; vocabSize includes ids above model.vocab
  • With the ModernBERT file (tokenizer/tokenizer.json) and merges loading (sibling lane), " leading" encodes to [50275, 16378] (reference: [' ', 'leading']) instead of [245, 4283, …], and vocabSize == 50368
  • CHANGELOG.md notes the vocabSize behaviour change
  • Result reported back on #1321

Notes

The flags single_word, lstrip, rstrip, normalized exist on every entry. In the two Laya files they are all false except normalized: true on the non-special ModernBERT entries (and lstrip: true on one special token). Implement plain longest-match as the existing special-token code does; if a file you care about sets single_word or lstrip/rstrip, open a separate issue rather than growing this one. #912 asks for the same vocabSize fix on the SentencePiece side — closing that part here is fine, say so in the PR.

Lenguaje dominante
Kotlin
Estrellas
52
Forks
15
Merge medio
1 d 15 h
PR fusionados (30 d)
36

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de SKaiNET-developers/SKaiNET

Todos los issues de SKaiNET-developers/SKaiNET

Issues similares

Más issues de Kotlin

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.