BertTokenizer merges words separated only by `\n`, `\t` or `\r`
Los mantenedores suelen responder en 2 días
@rosebyte ya está trabajando en esto.
Desde el 23/9/2026.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 85/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- csharp
- Área
- machine-learning
Línea de trabajo
Comienza en src/Microsoft.ML.Tokenizers/Normalizer/BertNormalizer.cs e inspecciona cómo se gestionan los caracteres de UnicodeCategory.Control antes de la pretokenización por espacios en blanco. Reproduce el problema con los ejemplos proporcionados one\nline, one\tline y note\nbook; se considera terminado cuando estos caracteres se tratan como espacios en blanco y las salidas coinciden con la tokenización esperada.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Description
BertTokenizer deletes \n, \t and \r rather than treating them as whitespace, so the words on either side are joined into one word before tokenization. What that yields depends on the vocabulary: one\nline becomes one ##line, while note\nbook becomes the single token notebook. Either way it changes the tokens, and therefore the embeddings, for any text with a line break or tab that has no adjacent space, which is common in real documents.
Hugging Face's BERT basic tokenizer treats these three characters as whitespace: _is_control explicitly excludes them, and _is_whitespace includes them. That is not specific to HF — it comes from the original BERT release, whose tokenization.py returns False from _is_control for \t, \n and \r, commented "These are technically control characters but we count them as whitespace characters". BertTokenizer's own remarks state that its implementation "is based on the original Bert implementation in the Hugging Face Transformers library".
If this is deliberate — a decision to treat Control uniformly rather than follow the _is_whitespace exception — please share the reasoning, and I'll close this.
Versions
- Microsoft.ML.Tokenizers 2.0.0 and 3.0.0-preview.26457.2, with identical results
- .NET 10.0.12, macOS arm64
- Vocabulary:
sentence-transformers/all-MiniLM-L6-v2vocab.txt(uncased BERT, 30,522 tokens) - Reference: Hugging Face
tokenizers0.21.1,BertWordPieceTokenizer(vocab, lowercase=True) - Cosine under "Impact":
onnxruntime1.22.0, mean pooling over the token ids
Reproduce
using Microsoft.ML.Tokenizers;
var vocab = File.ReadAllLines("vocab.txt");
var tokenizer = BertTokenizer.Create("vocab.txt", new BertOptions());
foreach (var input in new[] { "one\nline", "one\tline", "one \nline", "note\nbook" })
{
var tokens = tokenizer.EncodeToIds(input).Select(id => vocab[id]);
Console.WriteLine($"{input.Replace("\n", "\\n").Replace("\t", "\\t")} -> {string.Join(' ', tokens)}");
}
Expected vs actual
| input | actual | expected (HF tokenizers) |
|---|---|---|
one\nline |
[CLS] one ##line [SEP] |
[CLS] one line [SEP] |
one\tline |
[CLS] one ##line [SEP] |
[CLS] one line [SEP] |
one \nline |
[CLS] one line [SEP] |
[CLS] one line [SEP] |
note\nbook |
[CLS] notebook [SEP] |
[CLS] note book [SEP] |
No BertOptions setting preserves BERT basic-tokenization semantics while correcting this behavior. Setting ApplyBasicTokenization = false bypasses the faulty normalizer, but also disables the rest of BERT basic tokenization.
Cause
BertNormalizer.Normalize (src/Microsoft.ML.Tokenizers/Normalizer/BertNormalizer.cs) skips every character whose category is UnicodeCategory.Control:
if (category == UnicodeCategory.Control)
{
i += inc;
continue;
}
\t, \n and \r are all Control, so they are removed before the pre-tokenizer splits on whitespace. A fix is to emit a space for those three characters before that check, as HF does:
if (c is '\t' or '\n' or '\r')
{
AddChar(ref buffer, ref index, ' ');
continue;
}
Impact
Embedding all-MiniLM-L6-v2 through ONNX Runtime with mean pooling, the input " line one\n\tline two " scores cosine 0.936 against the same text tokenized by HF tokenizers. Callers can work around it by replacing \t, \n and \r with spaces before encoding.
- Lenguaje dominante
- C#
- Estrellas
- 9.4k
- Forks
- 2k
- Merge medio
- 4 d 2 h
- PR fusionados (30 d)
- 20
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de dotnet/machinelearning
-
LlamaTokenizer encodes some runs of repeated characters differently from SentencePiecePosiblemente ocupada @Laurianti la tomó hace 9 días. Abiertountriaged
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
dotnet/machinelearning#7751 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Comment in example code not matching docstring for TimeSeriesCatalog.DetectSeasonality MethodAbiertoarea-TimeSeries documentation untriaged
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
dotnet/machinelearning#7205 ·
Los mantenedores suelen responder en 2 días
-
area-ONNX documentation Priority:3
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
dotnet/machinelearning#5894 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Add a .vsconfig fileAbiertoarea-Infrastructure Priority:2
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
dotnet/machinelearning#3260 · 2 comentarios · 1 reacción ·
Los mantenedores suelen responder en 2 días
-
area-Core enhancement Priority:3
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
dotnet/machinelearning#2669 · 1 comentario ·
Los mantenedores suelen responder en 2 días
Todos los issues de dotnet/machinelearning
Issues similares
-
VideoViewer: rotated (portrait phone) videos shown sideways when system decimal separator is a commaAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Volodymyr-Petrunin/Bankomaten#45 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
AvaloniaUI/Avalonia#22420 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
dotnet/SqlClient#4823 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
type/automation type/tech-debt
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día