Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Segfault: 4+ unchecked `PyBytes_AsStringAndSize` on user `read()` return uses uninitialized memory

Abierto
#292 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 2 días

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-2 días
Aptitud para principiantes
76/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Tranquilo
Stack tecnológico
c, python
Área
backend

Línea de trabajo

Comienza en c-ext/compressor.c, c-ext/decompressor.c, c-ext/compressoriterator.c y c-ext/decompressoriterator.c, en las llamadas a PyBytes_AsStringAndSize indicadas; después, inspecciona las llamadas adicionales en read_compressor_input y su equivalente para el descompresor. Reproduce el problema con los ejemplos de BadSource proporcionados y verifica que las devoluciones que no sean bytes generen TypeError sin provocar un fallo ni una interrupción por aserción en ninguna de las rutas de streaming afectadas.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Summary

Several streaming paths call PyBytes_AsStringAndSize(result, &readBuffer, &readSize) on the return value of a caller-supplied source.read() and do not check the return value. When read() returns a non-bytes object (e.g., str, None, bytearray), readBuffer and readSize are left with their prior/uninitialized contents; the next ZSTD_compressStream2 call reads from those addresses. Observed: SEGV on release builds, _Py_CheckFunctionResult abort on debug builds.

Impact

  • Severity: SEGV on release builds; assertion abort on debug builds.
  • Reachability: Any caller-supplied source.read() whose return is not bytes. A trivial wrapper around a text file or a mistakenly-returned bytearray triggers it.
  • Version: 0.25.0 (commit 7a77a75).
  • Platform: Confirmed Linux x86_64 / CPython 3.14 debug; bug is platform-independent.

Reproducers

SEGV on release — via copy_stream:

import zstandard, io

class BadSource:
    def read(self, size):
        return 'not bytes'   # str, not bytes

comp = zstandard.ZstdCompressor()
comp.copy_stream(BadSource(), io.BytesIO())
# Segmentation fault

Assertion abort on debug — via iterator:

import zstandard

class BadSource:
    def read(self, size):
        return 'not bytes'

comp = zstandard.ZstdCompressor()
it = comp.read_to_iter(BadSource())
next(it)
# Fatal Python error: _Py_CheckFunctionResult: a function returned a result with an exception set
# TypeError: expected bytes, str found

Root cause

PyBytes_AsStringAndSize returns -1 and sets a TypeError when its argument is not a bytes object. On failure, the by-address output parameters readBuffer / readSize are not written. zstandard ignores the return code and proceeds to use those addresses, passing them to ZSTD_compressStream2 which reads from whatever happens to be on the stack / in registers.

Affected sites

File Line Function
c-ext/compressor.c 349 copy_stream
c-ext/decompressor.c 202 decompressor_copy_stream
c-ext/compressoriterator.c 98 ZstdCompressorIterator_iternext
c-ext/decompressoriterator.c 134 ZstdDecompressorIterator_iternext

Plus 2 additional sites reported in the full analysis (in read_compressor_input and the decompressor equivalent).

Suggested fix

Mechanical — add the standard error check after every call:

if (PyBytes_AsStringAndSize(result, &readBuffer, &readSize) < 0) {
    Py_DECREF(result);
    goto finally;
}

Optionally: tighten the documented API contract on source.read() to specify that the return must be a bytes object (the C code already expects this). Enforcing it at the boundary would be a small additional cleanup.

Methodology

Found via cext-review-toolkit (Tree-sitter-based static analysis with structured naive/informed review passes). SEGV on release verified at the copy_stream site; assertion abort on debug verified at the iterator site. Four sites confirmed via direct reproducer; two more confirmed via static review. Happy to open a PR — the fix is a ~8-line diff.

Discovery, root-cause analysis, and issue drafting were performed by Claude Code and reviewed by a human before filing.

Full report

Complete multi-agent analysis (48 FIX findings across 13 categories, plus a reproducer appendix): https://gist.github.com/devdanzin/b86039ac097141579590c1a0f3a43605

Lenguaje dominante
C
Estrellas
641
Forks
117
Merge medio
1 d 14 h
PR fusionados (30 d)
5

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de indygreg/python-zstandard

Todos los issues de indygreg/python-zstandard

Issues similares

Más issues de C

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.