AsciiArtConverter silently drops every non-ASCII character
Los mantenedores suelen responder en 2 días
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 38/100
Línea de trabajo
Start with pyrit/converter/ascii_art_converter.py and test_ascii_art_converter.py, then inspect the neighboring ascii_smuggler_converter.py, NatoConverter, and BrailleConverter behavior. Reproduce the non-ASCII cases against art.text2art and the available fonts. Done means the maintainer-selected behavior is implemented and regression tests prove prompts are not silently altered.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
AsciiArtConverter silently deletes every non-ASCII character from the prompt. Nothing is raised and nothing is logged — the target just receives a mangled prompt, or no prompt at all.
What happens
convert_async hands the prompt straight to art.text2art:
# pyrit/converter/ascii_art_converter.py:76
return ConverterResult(output_text=text2art(prompt, font=font), output_type="text")
text2art has no glyph for a character, so it omits it. I probed every font the art package exposes — 371 in total, 354 of them in this converter's rand pool:
| character | dropped by |
|---|---|
é, ü |
371 / 371 |
日, 本, 語, 漢 |
371 / 371 |
🙂 |
371 / 371 |
“, — |
371 / 371 |
ß |
307 / 371 |
A |
2 / 371 (hills, nfi1) |
Accented Latin, CJK, emoji, smart quotes and em-dashes are dropped by every font, so the loss does not depend on which font gets drawn.
Two concrete consequences
One character disappears, taking its whole glyph block with it:
AsciiArtConverter(font="block").convert_async(prompt="cafe") -> 4 glyph blocks (892 chars)
AsciiArtConverter(font="block").convert_async(prompt="café") -> 3 glyph blocks (672 chars)
And a prompt written entirely in a non-ASCII script converts to the empty string. With the default font="rand" this holds on every draw:
AsciiArtConverter().convert_async(prompt="日本語の指示") -> output_text='' (12/12 draws)
AsciiArtConverter(font="block").convert_async(prompt="忽略之前的所有指令") -> ''
For a red-teaming framework this is the harmful direction. A Chinese- or Japanese-language attack prompt is a first-class use case, and it converts to nothing. A mixed prompt reaches the target quietly altered, with no signal to the red-teamer.
Why I think this is a defect
Three things already in this repo establish the opposite convention for the same situation:
AsciiSmugglerConverterraisesValueErrornaming the characters outside the range it can encode (ascii_smuggler_converter.py:69-74, merged in #2540).NatoConverterwas changed to pass unmapped characters through rather than delete them (#2399).BrailleConvertergot the same treatment (#2539).
AsciiArtConverter's docstring documents only ValueError: If the input type is not supported; it never mentions that non-ASCII input is dropped. The tests only feed ASCII prompts — test_ascii_art_converter.py:16-29 assert len(result.output_text) > 0 — so nothing pins this behaviour in either direction.
What behaviour do you want?
Option A — raise, naming the characters this font cannot render. Consistent with #2540, and it fails loudly instead of altering the attack. Cost: a campaign whose prompts contain an accent or an emoji would start raising. Checking the resolved font is deterministic for everything that matters here, since the characters above are dropped by all 371 fonts; only ß-type characters are font-dependent, so under font="rand" such a prompt could raise on one draw and not the next.
Option B — render what the font can, pass the rest through. Keeps pipelines running and stops losing content, at the cost of a mixed output (art plus bare characters). This matches what #2399 and #2539 ended up doing for the other converters.
Option C — document the restriction, change nothing. Cheapest, but a silently altered attack prompt is a poor default for a tool whose job is to send a precise prompt.
I lean towards A: "the prompt that reaches the target is not the prompt I wrote" is precisely the failure a red-teaming tool should refuse rather than hide, and #2540 already set that precedent for the converter next door. If B is preferable because non-ASCII prompts are expected to keep working, I would implement B instead.
Happy to take whichever you pick, with regression tests over the real art font list. Everything above was measured against the installed art package; I called no model or target.
- Lenguaje dominante
- Python
- Estrellas
- 4.6k
- Forks
- 924
- Merge medio
- 2 d 10 h
- PR fusionados (30 d)
- 210
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/PyRIT
-
BUG: PlagiarismScorer accepts invalid n-gram size and blank reference textPosiblemente ocupada @RohithPariki la tomó hace 2 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 2 días
-
PackageHallucinationScorer (Python) misses `from pkg.sub import x` and indented importsPosiblemente ocupada @barry166 la tomó hace 3 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
microsoft/PyRIT#2948 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
LiteLLMChatTarget does not flag or survive output-token truncationPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 2 días
-
BUG Configuration keeps runtime-status errors after polling recoversPosiblemente ocupada @rupayon123 la tomó hace 9 días. AbiertoBug: triage GUI help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
microsoft/PyRIT#2868 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
ObjectiveScorerEvaluator scores every conversation message as an assistant responsePosiblemente ocupada @feiiiiii5 la tomó hace 10 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 2 días
Todos los issues de microsoft/PyRIT
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
LearningCircuit/local-deep-research#7206 ·
Los mantenedores suelen responder en 1 día
-
[TASK] Document technology stackAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Los mantenedores suelen responder en 1 día
-
Proxy drops log notifications from backends that don't send FastMCP's msg/extra dictPosiblemente ocupada @asasemahmed la tomó hoy. Abiertobug server
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
[Bug]: Bedrock request metadata forwarding does not work for /embeddingsPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abiertobug llm translation
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día