Segfault: 4+ unchecked `PyBytes_AsStringAndSize` on user `read()` return uses uninitialized memory
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 76/100
Línea de trabajo
Comienza en c-ext/compressor.c, c-ext/decompressor.c, c-ext/compressoriterator.c y c-ext/decompressoriterator.c, en las llamadas a PyBytes_AsStringAndSize indicadas; después, inspecciona las llamadas adicionales en read_compressor_input y su equivalente para el descompresor. Reproduce el problema con los ejemplos de BadSource proporcionados y verifica que las devoluciones que no sean bytes generen TypeError sin provocar un fallo ni una interrupción por aserción en ninguna de las rutas de streaming afectadas.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
Several streaming paths call PyBytes_AsStringAndSize(result, &readBuffer, &readSize) on the return value of a caller-supplied source.read() and do not check the return value. When read() returns a non-bytes object (e.g., str, None, bytearray), readBuffer and readSize are left with their prior/uninitialized contents; the next ZSTD_compressStream2 call reads from those addresses. Observed: SEGV on release builds, _Py_CheckFunctionResult abort on debug builds.
Impact
- Severity: SEGV on release builds; assertion abort on debug builds.
- Reachability: Any caller-supplied
source.read()whose return is notbytes. A trivial wrapper around a text file or a mistakenly-returnedbytearraytriggers it. - Version: 0.25.0 (commit
7a77a75). - Platform: Confirmed Linux x86_64 / CPython 3.14 debug; bug is platform-independent.
Reproducers
SEGV on release — via copy_stream:
import zstandard, io
class BadSource:
def read(self, size):
return 'not bytes' # str, not bytes
comp = zstandard.ZstdCompressor()
comp.copy_stream(BadSource(), io.BytesIO())
# Segmentation fault
Assertion abort on debug — via iterator:
import zstandard
class BadSource:
def read(self, size):
return 'not bytes'
comp = zstandard.ZstdCompressor()
it = comp.read_to_iter(BadSource())
next(it)
# Fatal Python error: _Py_CheckFunctionResult: a function returned a result with an exception set
# TypeError: expected bytes, str found
Root cause
PyBytes_AsStringAndSize returns -1 and sets a TypeError when its argument is not a bytes object. On failure, the by-address output parameters readBuffer / readSize are not written. zstandard ignores the return code and proceeds to use those addresses, passing them to ZSTD_compressStream2 which reads from whatever happens to be on the stack / in registers.
Affected sites
| File | Line | Function |
|---|---|---|
c-ext/compressor.c |
349 | copy_stream |
c-ext/decompressor.c |
202 | decompressor_copy_stream |
c-ext/compressoriterator.c |
98 | ZstdCompressorIterator_iternext |
c-ext/decompressoriterator.c |
134 | ZstdDecompressorIterator_iternext |
Plus 2 additional sites reported in the full analysis (in read_compressor_input and the decompressor equivalent).
Suggested fix
Mechanical — add the standard error check after every call:
if (PyBytes_AsStringAndSize(result, &readBuffer, &readSize) < 0) {
Py_DECREF(result);
goto finally;
}
Optionally: tighten the documented API contract on source.read() to specify that the return must be a bytes object (the C code already expects this). Enforcing it at the boundary would be a small additional cleanup.
Methodology
Found via cext-review-toolkit (Tree-sitter-based static analysis with structured naive/informed review passes). SEGV on release verified at the copy_stream site; assertion abort on debug verified at the iterator site. Four sites confirmed via direct reproducer; two more confirmed via static review. Happy to open a PR — the fix is a ~8-line diff.
Discovery, root-cause analysis, and issue drafting were performed by Claude Code and reviewed by a human before filing.
Full report
Complete multi-agent analysis (48 FIX findings across 13 categories, plus a reproducer appendix): https://gist.github.com/devdanzin/b86039ac097141579590c1a0f3a43605
- Lenguaje dominante
- C
- Estrellas
- 641
- Forks
- 117
- Merge medio
- 1 d 14 h
- PR fusionados (30 d)
- 5
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de indygreg/python-zstandard
-
`multi_decompress_to_buffer([])` terminates the process with SIGFPEPosiblemente ocupada @mikamikasuki la tomó hace 6 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
indygreg/python-zstandard#335 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
indygreg/python-zstandard#295 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 3/5 1-2 días Aptitud para principiantes 64/100
indygreg/python-zstandard#345 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 55/100
indygreg/python-zstandard#334 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 3/5 1-2 días Aptitud para principiantes 54/100
indygreg/python-zstandard#333 ·
Los mantenedores suelen responder en 2 días
Todos los issues de indygreg/python-zstandard
Issues similares
-
Policy query leaks host primary block (BSL_PrimaryBlock_deinit skipped) on two early-exit pathsAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
NASA-AMMOS/BSL#355 ·
Los mantenedores suelen responder en 1 día
-
#242 leftovers: dated narrative and shas in the social-features test planPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
EchoTools/nevr-runtime#264 ·
Los mantenedores suelen responder en 1 día
-
area/docdb kind/bug priority/medium status/awaiting-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
yugabyte/yugabyte-db#34873 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
[sqlcipher] update to 4.19.0Abiertocategory:port-update
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 2 días