sparklemotion/nokogiri

JRuby: UtfHelpper.writeCharToUtf8 cannot handle unicode supplementary character

Ouverte

#2 410 ouverte le 5 janv. 2022

 (9 commentaires) (0 réaction) (0 personne assignée)Ruby (806 forks)batch import
blockedhelp wantedplatform/jrubytopic/encoding

Métriques du dépôt

Stars
 (5 615 étoiles)
Métriques de merge PR
 (Merge moyen 1j 7h) (14 PRs mergées en 30 j)

Description

https://github.com/sparklemotion/nokogiri/blob/55029bfba481338825c99e78af2b182b1cc49e04/ext/java/nokogiri/internals/c14n/UtfHelpper.java#L51

since the Canonicalizer process input String character by character. Java uses 16 bits to represent a character; when the input string contains Unicode characters whose code pen are larger than 0Xffff(65535) it will be split into two char, since neither char will not be
recognized, the Unicode characters will be transferred to 2 ??(3f) instead.

for example, if I want to canonicalize an input that contains 𡏅 via c14n, in the output, 𡏅 will be replaced with ??

Guide contributeur