Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

BertTokenizer merges words separated only by `\n`, `\t` or `\r`

Abierto Apto para principiantes
#7,724 2 comentarios 0 reacciones 1 asignado Ver en GitHub

Los mantenedores suelen responder en 2 días

@rosebyte ya está trabajando en esto.

Desde el 23/9/2026.

Evaluación

Dificultad
2/5
Tiempo estimado
1-3 horas
Aptitud para principiantes
85/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
csharp

Línea de trabajo

Comienza en src/Microsoft.ML.Tokenizers/Normalizer/BertNormalizer.cs e inspecciona cómo se gestionan los caracteres de UnicodeCategory.Control antes de la pretokenización por espacios en blanco. Reproduce el problema con los ejemplos proporcionados one\nline, one\tline y note\nbook; se considera terminado cuando estos caracteres se tratan como espacios en blanco y las salidas coinciden con la tokenización esperada.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

area-Tokenizers green

Description

BertTokenizer deletes \n, \t and \r rather than treating them as whitespace, so the words on either side are joined into one word before tokenization. What that yields depends on the vocabulary: one\nline becomes one ##line, while note\nbook becomes the single token notebook. Either way it changes the tokens, and therefore the embeddings, for any text with a line break or tab that has no adjacent space, which is common in real documents.

Hugging Face's BERT basic tokenizer treats these three characters as whitespace: _is_control explicitly excludes them, and _is_whitespace includes them. That is not specific to HF — it comes from the original BERT release, whose tokenization.py returns False from _is_control for \t, \n and \r, commented "These are technically control characters but we count them as whitespace characters". BertTokenizer's own remarks state that its implementation "is based on the original Bert implementation in the Hugging Face Transformers library".

If this is deliberate — a decision to treat Control uniformly rather than follow the _is_whitespace exception — please share the reasoning, and I'll close this.

Versions

  • Microsoft.ML.Tokenizers 2.0.0 and 3.0.0-preview.26457.2, with identical results
  • .NET 10.0.12, macOS arm64
  • Vocabulary: sentence-transformers/all-MiniLM-L6-v2 vocab.txt (uncased BERT, 30,522 tokens)
  • Reference: Hugging Face tokenizers 0.21.1, BertWordPieceTokenizer(vocab, lowercase=True)
  • Cosine under "Impact": onnxruntime 1.22.0, mean pooling over the token ids

Reproduce

using Microsoft.ML.Tokenizers;

var vocab = File.ReadAllLines("vocab.txt");
var tokenizer = BertTokenizer.Create("vocab.txt", new BertOptions());
foreach (var input in new[] { "one\nline", "one\tline", "one \nline", "note\nbook" })
{
    var tokens = tokenizer.EncodeToIds(input).Select(id => vocab[id]);
    Console.WriteLine($"{input.Replace("\n", "\\n").Replace("\t", "\\t")} -> {string.Join(' ', tokens)}");
}

Expected vs actual

input actual expected (HF tokenizers)
one\nline [CLS] one ##line [SEP] [CLS] one line [SEP]
one\tline [CLS] one ##line [SEP] [CLS] one line [SEP]
one \nline [CLS] one line [SEP] [CLS] one line [SEP]
note\nbook [CLS] notebook [SEP] [CLS] note book [SEP]

No BertOptions setting preserves BERT basic-tokenization semantics while correcting this behavior. Setting ApplyBasicTokenization = false bypasses the faulty normalizer, but also disables the rest of BERT basic tokenization.

Cause

BertNormalizer.Normalize (src/Microsoft.ML.Tokenizers/Normalizer/BertNormalizer.cs) skips every character whose category is UnicodeCategory.Control:

if (category == UnicodeCategory.Control)
{
    i += inc;
    continue;
}

\t, \n and \r are all Control, so they are removed before the pre-tokenizer splits on whitespace. A fix is to emit a space for those three characters before that check, as HF does:

if (c is '\t' or '\n' or '\r')
{
    AddChar(ref buffer, ref index, ' ');
    continue;
}

Impact

Embedding all-MiniLM-L6-v2 through ONNX Runtime with mean pooling, the input " line one\n\tline two " scores cosine 0.936 against the same text tokenized by HF tokenizers. Callers can work around it by replacing \t, \n and \r with spaces before encoding.

Lenguaje dominante
C#
Estrellas
9.4k
Forks
2k
Merge medio
4 d 2 h
PR fusionados (30 d)
20

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de dotnet/machinelearning

Todos los issues de dotnet/machinelearning

Issues similares

Más issues de C#

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.