Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Lane 2 · kotlin-core] tokenizer.json parity — apply the normalizer block (Replace, Prepend, Sequence, NFC) before pre-tokenization

Abierto
#1,325 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
1-2 días
Aptitud para principiantes
43/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
kotlin
Área
backend, data

Línea de trabajo

Start with SpecialTokenSplitter.kt and the normalizer handling in SentencePieceTokenizer.kt#L382-L391; the issue names both tokenizer factories to update and the existing source sets for platform NFC code. First check the Lane 1 survey for the NFC strategy, then implement the normalizer plumbing while it is open. Run the listed jvmTest and linuxX64Test; done means the three specified normalizers pass, unknown types fail with their names, and the ModernBERT fixture ends with id 3974.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

coding size:m skill:kotlin-core sub-issue

Sub-issue of #1321 (HF tokenizer.json parity for BPE encoders).

Lane: 2 · Kotlin core
Skill needed: Kotlin Multiplatform (expect/actual or a pure-Kotlin table), Unicode normalization basics. No tensor knowledge.
Size: m (1–2 days; most of it is the NFC implementation and its tests)
Blocked by: the Lane 1 survey (it decides the NFC strategy). Start with the Replace/Sequence plumbing while it is open.

What to do

  1. Add a small Normalizer step to sk.ainet.io.tokenizer that is built from tokenizer.json#normalizer and applied to each non-special segment before pre-tokenization (the reference applies it before added-token matching only for tokens with normalized: true; applying it to the segments between atomic matches is the observable behaviour for the files in the parent).
  2. Support, in this order: null (no-op), Replace (pattern.String → content; pattern.Regex can throw UnsupportedTokenizerException with the pattern in the message), Prepend, Sequence (compose), NFC. Everything else: throw UnsupportedTokenizerException naming the type — do not silently skip, that is the current bug.
  3. Implement NFC the way Lane 1 recommended. If that is expect/actual: JVM via java.text.Normalizer, JS/Wasm via normalize("NFC"), Kotlin/Native passthrough with a documented warning in the KDoc. Put the per-target code in the existing jvmMain / jsMain / wasmJsMain / native64Main source sets of skainet-io-core.
  4. Wire it into QwenByteLevelBpeTokenizer.fromTokenizerJson and SentencePieceTokenizer.fromTokenizerJson; the existing Prepend sniffing at SentencePieceTokenizer.kt#L382-L391 should become a consumer of the new step rather than a second parser.
  5. Tests in commonTest: Replace, Sequence, and NFC on "e\u0301" → "é" (JVM at least; mark the native expectation explicitly). Run ./gradlew :skainet-io:skainet-io-core:jvmTest :skainet-io:skainet-io-core:linuxX64Test.

Acceptance

  • ModernBERT file: the third fixture text of #1321 (ending in decomposed e + U+0301) encodes identically to the reference — last id 3974, not 299, 136, 212
  • Unknown normalizer types fail loudly with the type name in the message
  • Tests green on jvmTest and linuxX64Test
  • Result reported back on #1321

Notes

Sibling file to copy the shape from: SpecialTokenSplitter.kt is a small decorator built by the factory from a JSON block — the normalizer is the same pattern one step earlier in the pipeline. Keep the implementation out of QwenByteLevelBpeTokenizer itself so the Metaspace lane can reuse it.

Lenguaje dominante
Kotlin
Estrellas
52
Forks
15
Merge medio
1 d 15 h
PR fusionados (30 d)
36

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de SKaiNET-developers/SKaiNET

Todos los issues de SKaiNET-developers/SKaiNET

Issues similares

Más issues de Kotlin

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.