`TextEncoder.encodeInto()` underfills the destination for some non-ASCII text
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 84/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- cpp, javascript, nodejs
Direzione di ricerca
Inizia in src/encoding_binding.cc, soprattutto con simpleUtfEncodingLength() e findBestFit(), poi esamina la copertura WPT in encodeInto.any.js. Riproduci gli esempi di 33 caratteri e aggiungi la copertura per U+0400–U+07FF e per le destinazioni strette; il lavoro è completato quando encodeInto riporta i valori attesi di lettura/scrittura e i test WPT pertinenti passano.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
There are two problems I found with TextEncoder.encodeInto(), both of which can cause encoding to stall even when the next char can fit the destination.
-
A 2-byte char requires a 3-byte destination
const encoder = new TextEncoder(); const text = '\u0400'.repeat(33); console.log(encoder.encodeInto(text, new Uint8Array(2))); // { read: 0, written: 0 } console.log(encoder.encodeInto(text, new Uint8Array(3))); // { read: 1, written: 2 }The second call proves that '\u0400' should fit into a 2-byte array.
-
Appending an unread character changes encoding progress
const encoder = new TextEncoder(); const text = 'é'.repeat(33); console.log(encoder.encodeInto(text, new Uint8Array(2))); // { read: 0, written: 0 } console.log(encoder.encodeInto(text + '☺', new Uint8Array(2))); // { read: 1, written: 2 }Appending
☺should not change whether preceding chars can be read into the buffer, but there it is.
The bugs were introduced by the encodeInto() performance change in Node.js v25.4.0. The examples above use length 33 strings to exercise that optimized path(kSmallStringThreshold = 32). Unfortunately the current encodeInto.any.js WPT tests fail to expose the problems because:
- all input cases use 7 or fewer code units
- even then, the cases don't use chars between U+0400 and U+07FF, and
- their cases don't contain a narrow dst capacity to reveal the signed-byte problem.
Proposed fixes
-
Incorrect cutoff in
simpleUtfEncodingLength()- if (c < 0x400) return 2; + if (c < 0x800) return 2;(very likely a typo, given the comment immediately below it says "Code points < 0x800: 2 bytes")
-
Signed-byte handling in
findBestFit()- size_t extra = simpleUtfEncodingLength(data[pos]); + size_t extra = simpleUtfEncodingLength(UTF16 ? data[pos] : static_cast<uint8_t>(data[pos]));
- Lingua principale
- JavaScript
- Stelle
- 122k
- Fork
- 37.4k
- Merge medio
- 4g 3h
- PR unite (30g)
- 279
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di nodejs/node
-
doc
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
-
build
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
-
feature request
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
stale
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
palladius/rails8-app-on-gcp#145 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
dotenvx/dotenv-vscode#139 ·
-
test-change-proposal
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
web-platform-tests/interop#1455 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
corsairdev/corsair#1764 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100