Memory leaks in stream lifecycle: 4 `__exit__` + 2 `close()` discard CallMethod returns; 5 types leak on `__init__` re-call (~5.5 KB); 2 `Py_buffer` leaks on closed streams
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 48/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 領域
- backend, performance
調査の方向性
c-ext/compressionreader.c、c-ext/compressionwriter.c、c-ext/decompressionwriter.c、c-ext/decompressionreader.c にある4つの exit/close 箇所から始め、続いて影響を受ける5つの init 実装と、issue で名前が挙げられている writer メソッドを調査します。提供されている tracemalloc、RSS、Py_buffer の再現プログラムを実行します。各リークに対処し、再初期化の動作を明示的に選択して検証できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Summary
Four distinct memory-leak patterns in stream/resource lifecycle. Three are mechanical (discarded CallMethod return values, missing PyBuffer_Release on a closed-stream branch); one is semantic (re-callable __init__ leaks the prior ZSTD contexts). Filing together because all four surface as "memory grows during normal usage of stream writers/readers", but each has a distinct fix and can be addressed independently.
Impact
- Severity: Memory leak — no crash. Magnitude per occurrence ranges from ~30 bytes (discarded
CallMethodreturns) to ~5.5 KB (re-init contexts). - Reachability: Standard idioms —
with comp.stream_writer(...):, explicit.close(),.write()after close. - Version: 0.25.0 (commit
7a77a75). - Platform: Confirmed Linux x86_64 / CPython 3.14 debug; bug is platform-independent.
Leak 1: 4 __exit__ methods discard close() return — ~31 B per with exit
PyObject_CallMethod(self, "close", NULL) returns a new reference; all 4 __exit__ implementations discard it without Py_DECREF.
Reproducer:
import zstandard, tracemalloc, gc
tracemalloc.start(); gc.collect()
s1 = tracemalloc.take_snapshot()
for _ in range(5000):
comp = zstandard.ZstdCompressor()
with comp.stream_writer(open('/dev/null', 'wb')) as w:
w.write(b'hello' * 100)
gc.collect()
s2 = tracemalloc.take_snapshot()
diff = sum(s.size_diff for s in s2.compare_to(s1, 'lineno') if s.size_diff > 0)
print(f"{diff/5000:.1f} bytes per __exit__") # ~31.1
Sites:
c-ext/compressionreader.c:57(compressionreader_exit)c-ext/compressionwriter.c:53(ZstdCompressionWriter_exit)c-ext/decompressionwriter.c:41(ZstdDecompressionWriter_exit)c-ext/decompressionreader.c:57(decompressionreader_exit)
Fix:
PyObject *result = PyObject_CallMethod(self, "close", NULL);
Py_XDECREF(result);
Leak 2: 2 close() methods discard flush() return — ~32 B per close
Same pattern as Leak 1, different method. close() calls self.flush() via PyObject_CallMethod and discards the return.
Reproducer:
import zstandard, tracemalloc, gc
tracemalloc.start(); gc.collect()
s1 = tracemalloc.take_snapshot()
for _ in range(5000):
comp = zstandard.ZstdCompressor()
w = comp.stream_writer(open('/dev/null', 'wb'))
w.write(b'hello' * 100)
w.close()
gc.collect()
s2 = tracemalloc.take_snapshot()
diff = sum(s.size_diff for s in s2.compare_to(s1, 'lineno') if s.size_diff > 0)
print(f"{diff/5000:.1f} bytes per close") # ~31.7
Sites:
c-ext/compressionwriter.c:219(ZstdCompressionWriter_close)c-ext/decompressionwriter.c:155(ZstdDecompressionWriter_close)
Fix: Py_XDECREF(result); after each PyObject_CallMethod(..., "flush", ...) call.
Leak 3: Re-callable __init__ leaks ZSTD contexts — ~5.5 KB per re-init
Calling comp.__init__(...) on an already-initialized instance allocates a new cctx and params (via ZSTD_createCCtx / ZSTD_createCCtxParams) without freeing the prior ones. Uses system malloc, not CPython's allocator — tracemalloc doesn't observe it; RSS grows.
Affected types: ZstdCompressor, ZstdDecompressor, ZstdCompressionDict, BufferWithSegments, BufferWithSegmentsCollection.
Reproducer:
import zstandard, resource, gc
gc.collect()
r1 = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
comp = zstandard.ZstdCompressor()
for _ in range(50000):
comp.__init__()
gc.collect()
r2 = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
print(f"{(r2-r1)*1024/50000:.0f} bytes per re-init") # ~5541
Fix options:
Option A — free prior state at top of tp_init
if (self->cctx) { ZSTD_freeCCtx(self->cctx); self->cctx = NULL; }
if (self->params) { ZSTD_freeCCtxParams(self->params); self->params = NULL; }
/* ... then allocate ... */
Option B — reject re-init
if (self->cctx) {
PyErr_SetString(PyExc_RuntimeError, "already initialized");
return -1;
}
Option B composes cleanly with a separately-reported __new__() fix (if __new__ allocates the context in tp_new, tp_init simplifies to argument parsing and runtime-configuration only, and re-init naturally becomes "error").
Leak 4: 2 Py_buffer leaks in writer methods on closed streams
The y* arg format acquires a Py_buffer on input data. The "stream is closed" check returns NULL before PyBuffer_Release → buffer stays locked → a later bytearray.extend() on the same data raises BufferError, and the underlying memory is held until the buffer-owning object itself is collected.
Reproducer:
import zstandard
comp = zstandard.ZstdCompressor()
writer = comp.stream_writer(open('/dev/null', 'wb'))
writer.write(b'hello')
writer.close()
data = bytearray(1000)
try:
writer.write(data) # Py_buffer acquired, error, never released
except ValueError:
pass
data.extend(b'x') # BufferError: Existing exports of data
Sites:
c-ext/compressionwriter.c:85(ZstdCompressionWriter_write— closed-check path)c-ext/decompressionwriter.c:73(ZstdDecompressionWriter_write— closed-check path)c-ext/decompressionwriter.c:103(same function —output.dstleak onwriter.write()raising)
Fix:
if (self->closed) {
PyBuffer_Release(&source);
PyErr_SetString(PyExc_ValueError, "stream is closed");
return NULL;
}
Suggested PR shape
Four independent patches; happy to bundle in one PR or split by leak. The Py_XDECREF-on-CallMethod fixes (Leaks 1 + 2) are trivial; Leak 3 is a semantic choice (free-and-reinit vs. reject-reinit); Leak 4 is mechanical.
Methodology
Found via cext-review-toolkit (Tree-sitter-based static analysis with structured naive/informed review passes). All four leaks verified live on CPython 3.14.3 debug build. Leaks 1 + 2 measured via tracemalloc (per-call deltas match the reference-count-size overhead exactly). Leak 3 measured via resource.getrusage(ru_maxrss) because the allocator is libc malloc, outside CPython's tracking. Leak 4 verified via the BufferError observable. Happy to open a PR.
Discovery, root-cause analysis, and issue drafting were performed by Claude Code and reviewed by a human before filing.
Full report
Complete multi-agent analysis (48 FIX findings across 13 categories, plus a reproducer appendix): https://gist.github.com/devdanzin/b86039ac097141579590c1a0f3a43605
- 主要言語
- C
- スター
- 641
- フォーク
- 117
- 平均マージ
- 1日 14時間
- マージ済み PR(30日)
- 5
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
indygreg/python-zstandard のほかの issue
-
`multi_decompress_to_buffer([])` terminates the process with SIGFPE対応中かも @mikamikasuki が 5 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
indygreg/python-zstandard#335 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
indygreg/python-zstandard#295 ·
メンテナーはふだん 2 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 64/100
indygreg/python-zstandard#345 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 55/100
indygreg/python-zstandard#334 ·
メンテナーはふだん 2 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 54/100
indygreg/python-zstandard#333 ·
メンテナーはふだん 2 日以内に返信
indygreg/python-zstandard の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
-
backend
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
BasedHardware/omi#20940 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
kovidgoyal/kitty#10625 ·
メンテナーはふだん 1 日以内に返信
-
Feature Status: Needs Triage
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
メンテナーはふだん 1 日以内に返信