[Lane 2 · kotlin-core] tokenizer.json parity — match non-special added_tokens atomically and count them in vocabSize
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 72/100
Línea de trabajo
Read the added-token loop in QwenByteLevelBpeTokenizer.kt and TokenizerFactory.kt, then inspect the related tests in TokenizerFactoryDispatchTest.kt. Add a synthetic case where a non-special token would otherwise split, and check that it encodes atomically and vocabSize includes its id. Run ./gradlew :skainet-io:skainet-io-core:jvmTest; done means the test passes and the changelog records the vocabSize change.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).
Lane: 2 · Kotlin core
Skill needed: Kotlin. Helpful to have read how tokenizers matches added tokens (one page of docs, linked below). No tensor knowledge.
Size: s (a few hours)
Blocked by: nothing (independent of the merges lane; both touch the same file, so rebase once)
What to do
- Read how the reference matches added tokens: every entry of
added_tokensis matched atomically, before pre-tokenization, regardless ofspecial. Thespecialflag only decides whether the token is dropped bydecode(skip_special_tokens=True). See theAddedVocabularysection of thetokenizersdocs: https://huggingface.co/docs/tokenizers/api/added-tokens - In QwenByteLevelBpeTokenizer.kt#L257-L267 the loop skips entries with
"special": false. Register all entries in the atomic-match map. Keep a separateSet<Int>of the ids that are marked special so a later change can implementskip_special_tokens; nothing else needs that set today. - Make the same change in TokenizerFactory.kt#L155-L169 for the Unigram path (
wrapSentencePieceWithSpecialsFromJson). - Added tokens may have ids above
model.vocab(ModernBERT:[UNK]=50280 … and the whitespace runs up to 50367). Extend thetokensarray sovocabSizeequals the referenceget_vocab_size(with_added_tokens=True)— 50368 for ModernBERT instead of today's 50280 — and sodecodecan return their text. - Tests, in
TokenizerFactoryDispatchTest.kt: a synthetic file with onespecial: falseadded token whose content would otherwise BPE-split; assert it encodes to its id and thatvocabSizecounts it. Run./gradlew :skainet-io:skainet-io-core:jvmTest.
Acceptance
- Synthetic test: non-special added token encodes atomically;
vocabSizeincludes ids abovemodel.vocab - With the ModernBERT file (tokenizer/tokenizer.json) and merges loading (sibling lane),
" leading"encodes to[50275, 16378](reference:[' ', 'leading']) instead of[245, 4283, …], andvocabSize == 50368 -
CHANGELOG.mdnotes thevocabSizebehaviour change - Result reported back on #1321
Notes
The flags single_word, lstrip, rstrip, normalized exist on every entry. In the two Laya files they are all false except normalized: true on the non-special ModernBERT entries (and lstrip: true on one special token). Implement plain longest-match as the existing special-token code does; if a file you care about sets single_word or lstrip/rstrip, open a separate issue rather than growing this one. #912 asks for the same vocabSize fix on the SentencePiece side — closing that part here is fine, say so in the PR.
- Lenguaje dominante
- Kotlin
- Estrellas
- 52
- Forks
- 15
- Merge medio
- 1 d 15 h
- PR fusionados (30 d)
- 36
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de SKaiNET-developers/SKaiNET
-
coding good first issue size:xs skill:kotlin-core sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
SKaiNET-developers/SKaiNET#1323 ·
Los mantenedores suelen responder en 1 día
-
coding good first issue platform size:xs skill:js sub-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
SKaiNET-developers/SKaiNET#1232 ·
Los mantenedores suelen responder en 1 día
-
skeep tracking
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
SKaiNET-developers/SKaiNET#1331 ·
Los mantenedores suelen responder en 1 día
-
assessment size:s skill:review sub-issue
Dificultad 4/5 1-2 días Aptitud para principiantes 18/100
SKaiNET-developers/SKaiNET#1330 ·
Los mantenedores suelen responder en 1 día
-
documentation good first issue size:s skill:docs sub-issue
Dificultad 3/5 1-2 días Aptitud para principiantes 78/100
SKaiNET-developers/SKaiNET#1329 ·
Los mantenedores suelen responder en 1 día
Todos los issues de SKaiNET-developers/SKaiNET
Issues similares
-
[Submission] 抖音火山版Abiertosubmit-adaption submit-adaption-pre
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
BetterAndroid/android-notification-icon-project#744 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 64/100
utopia-rise/godot-jvm#1004 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
pedroSG94/RootEncoder#2213 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
[Feature]: AC charger voltageAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
dzid26/TeslaBatteryBLE#183 ·
Los mantenedores suelen responder en 1 día