Silent data-correctness bug: `readinto()` / `readinto1()` on `stream_reader` return `tell() == 0` after successful reads
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
调研方向
从 c-ext/compressionreader.c 中的 readinto 和 readinto1 方法体开始,将它们的本地 ZSTD_outBuffer 与 read() 路径中持久的输出记录进行比较。运行提供的 Python 复现程序,并验证 readinto() 和 readinto1() 之后的 tell() 是否与 read() 一致地递增;如果相关,请检查解压缩侧的模式。
由索引模型根据 Issue 内容生成。
描述
Summary
ZstdCompressionReader.readinto(buf) and .readinto1(buf) correctly copy compressed bytes into the caller's buffer, but the reader's internal bytesCompressed counter is not updated. As a result, stream_reader.tell() returns 0 after any number of successful readinto() calls, even though bytes were in fact written. The equivalent .read() path correctly advances the counter.
Impact
- Severity: Silent data-correctness bug — no crash, no exception, just a wrong value from
tell(). Any caller that relies ontell()to measure progress, compute offsets, or compare positions will silently malfunction. - Reachability: Standard
io.RawIOBaseidioms — any use ofreadinto()/readinto1()on astream_reader. Common in performance-sensitive decode pipelines that reuse a pre-allocated buffer. - Version: 0.25.0 (commit
7a77a75). - Platform: Platform-independent.
Reproducer
import zstandard, io
data = b'hello world ' * 10000
# read() — tell() works correctly
comp1 = zstandard.ZstdCompressor()
r1 = comp1.stream_reader(io.BytesIO(data))
r1.__enter__()
while r1.read(1024):
pass
print("read tell:", r1.tell()) # 29 (correct: total compressed bytes)
r1.__exit__(None, None, None)
# readinto() — tell() stuck at 0
comp2 = zstandard.ZstdCompressor()
r2 = comp2.stream_reader(io.BytesIO(data))
r2.__enter__()
buf = bytearray(1024)
while r2.readinto(buf):
pass
print("readinto tell:", r2.tell()) # 0 — BUG (should match read() path)
r2.__exit__(None, None, None)
Root cause
readinto / readinto1 build a ZSTD_outBuffer on the stack that wraps the caller's buffer:
ZSTD_outBuffer output = {dest, dest_size, 0};
zresult = ZSTD_compressStream2(cctx, &output, &input, ZSTD_e_continue);
After the call, output.pos holds the number of bytes written to dest. The reader then updates its position by reading from self->output.pos — the persistent struct, which the local-struct call never touched. So self->output.pos stays at zero, and bytesCompressed never advances.
The read() path uses self->output directly (not a local copy), so the persistent field is updated by ZSTD_compressStream2. That's why tell() works after read() but not after readinto.
Affected sites
c-ext/compressionreader.c—readintoandreadinto1method bodies.
(Same pattern may warrant a look on the decompression side as well, although the main analysis only flagged compression.)
Suggested fix
Two options; either is minimal.
Option A — advance from the local struct
Read bytesCompressed from the local output.pos before it goes out of scope:
ZSTD_outBuffer output = {dest, dest_size, 0};
zresult = ZSTD_compressStream2(cctx, &output, &input, ZSTD_e_continue);
/* ... */
self->bytesCompressed += output.pos; /* was: effectively + self->output.pos, i.e. 0 */
Option B — share the persistent struct
If you'd rather the two methods share the read() bookkeeping path, use self->output directly instead of a stack-local struct, and update self->output.dst / self->output.size to point at the caller's buffer before the call. Slightly more invasive but avoids duplicating the position-update logic.
Methodology
Found via cext-review-toolkit (Tree-sitter-based static analysis with structured naive/informed review passes). Reproducer verified live on CPython 3.14.3 debug build — read() path returns tell() == 29 (matches the compressed-output length); readinto() path returns tell() == 0 after the exact same compressed sequence is consumed. Happy to open a PR.
Discovery, root-cause analysis, and issue drafting were performed by Claude Code and reviewed by a human before filing.
Full report
Complete multi-agent analysis (48 FIX findings across 13 categories, plus a reproducer appendix): https://gist.github.com/devdanzin/b86039ac097141579590c1a0f3a43605
- 主要语言
- C
- 星标
- 641
- 派生
- 117
- 平均合并
- 1 天 14 小时
- 30 天内合并 PR
- 5
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
indygreg/python-zstandard 的其他 Issue
-
`multi_decompress_to_buffer([])` terminates the process with SIGFPE可能已有人在做 @mikamikasuki 于 5 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 78/100
indygreg/python-zstandard#335 ·
维护者通常 2 天内回复
-
难度 3/5 1-2 天 新手友好度 64/100
indygreg/python-zstandard#345 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 55/100
indygreg/python-zstandard#334 ·
维护者通常 2 天内回复
-
难度 3/5 1-2 天 新手友好度 54/100
indygreg/python-zstandard#333 ·
维护者通常 2 天内回复
-
难度 3/5 1-2 天 新手友好度 72/100
indygreg/python-zstandard#332 ·
维护者通常 2 天内回复
查看 indygreg/python-zstandard 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 78/100
-
backend
难度 2/5 1-3 小时 新手友好度 68/100
BasedHardware/omi#20940 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
kovidgoyal/kitty#10625 ·
维护者通常 1 天内回复
-
Feature Status: Needs Triage
难度 2/5 1-3 小时 新手友好度 73/100
维护者通常 1 天内回复