Byron/pulldown-cmark-to-cmark

Simplify backslash escapes where possible

Ouverte

#60 ouverte le 25 oct. 2023

 (2 commentaires) (0 réaction) (0 personne assignée)Rust (46 forks)auto 404
enhancementhelp wanted

Métriques du dépôt

Stars
 (60 étoiles)
Métriques de merge PR
 (Merge moyen 35m) (1 PR mergée en 30 j)

Description

Hi @Byron!

An issue was opened in the mdbook-i18n-helpers crate about how we treat backslashes in the translations: https://github.com/google/mdbook-i18n-helpers/issues/105.

As you might recall, the tooling there works by

  • parsing Markdown text -> Markdown AST
  • find translatable text in the AST
  • turn the AST nodes into Markdown text (using this crate)
  • translate this text and turn it into Markdown AST again

The third step here turns a Markdown file with

\x

into

\\x

This is completely valid! According to Backslash escapes, \\x and \x both mean backslash-x (2 bytes).

However, I can see how it could be confusing to people and tools which rely on a lot of backslashes, e.g., for LaTeX math like $\sqrt{\frac{1}{x}}$. Here, the translator will end up seeing the escaped backslashes: $\\sqrt{\\frac{1}{x}}$ because that is what we get back when we serialize the Markdown AST into Markdown text. It would be easier to work with the unescaped backslashes in this case.

So I'm proposint that pulldown-cmark-to-cmark would emit the simplest escaped form for an escaped character.

Guide contributeur