zlib: createZstdCompress appends an empty frame when end() is called with writes still queued
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 68/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- cpp, javascript, node.js
- Lĩnh vực
- backend, performance
Hướng nghiên cứu
Bắt đầu với bản tái hiện JavaScript được cung cấp, sau đó đọc lib/zlib.js, đặc biệt là ZlibBase#_transform và _flush. Kiểm tra ZstdCompressContext::DoThreadPoolWork trong src/node_zlib.cc và so sánh đường dẫn ghi được đưa vào hàng đợi với hành vi của gzip và brotli. Được xem là hoàn tất khi createZstdCompress phát ra một frame duy nhất với các byte giống hệt nhau cho một hoặc nhiều lần ghi và không có frame rỗng ở cuối.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Version
v26.9.0 (also v24.15.0)
Platform
Darwin 25.6.0 arm64 (also seen on Linux x64)
Subsystem
zlib
What steps will reproduce the bug?
'use strict';
const zlib = require('node:zlib');
function compress(create, writes) {
return new Promise((resolve, reject) => {
const stream = create();
const chunks = [];
stream.on('data', (chunk) => chunks.push(chunk));
stream.on('end', () => resolve(Buffer.concat(chunks)));
stream.on('error', reject);
for (const w of writes) stream.write(w);
stream.end();
});
}
(async () => {
const oneWrite = await compress(zlib.createZstdCompress, ['hello world']);
const twoWrites = await compress(zlib.createZstdCompress, ['hello ', 'world']);
console.log('one write: ', oneWrite.toString('hex'));
console.log('two writes:', twoWrites.toString('hex'));
console.log('extra bytes:', twoWrites.subarray(oneWrite.length).toString('hex'));
// Same input, queued writes: gzip and brotli are unaffected.
for (const [name, create, decompress] of [
['gzip', zlib.createGzip, zlib.gunzipSync],
['brotli', zlib.createBrotliCompress, zlib.brotliDecompressSync],
]) {
const a = await compress(create, ['hello world']);
const b = await compress(create, ['hello ', 'world']);
console.log(name, 'lengths', a.length, b.length, decompress(b).toString());
}
})();
How often does it reproduce? Is there a required condition?
Every time end() is called while at least one write is still queued in the compressor. In practice that is most pipelines whose source ends quickly, e.g. pipeline(tar.create(...), zlib.createZstdCompress(), fs.createWriteStream(...)): every archive we packed that way had the extra frame.
What is the expected behavior? Why is that the expected behavior?
One zstd frame, the same bytes whether the input arrived in one write or several:
one write: 28b52ffd005859000068656c6c6f20776f726c64
two writes: 28b52ffd005859000068656c6c6f20776f726c64
That's what gzip and brotli do. How the input was split into writes shouldn't change the compressed output.
What do you see instead?
one write: 28b52ffd005859000068656c6c6f20776f726c64
two writes: 28b52ffd005859000068656c6c6f20776f726c6428b52ffd2000010000
extra bytes: 28b52ffd2000010000
gzip lengths 31 31 hello world
brotli lengths 15 15 hello world
A second, empty zstd frame (28b52ffd 20 00 01 00 00: magic, single-segment descriptor with content size 0, one empty raw last block) is appended.
Additional information
I think the cause is in lib/zlib.js:
- In
ZlibBase#_transform, whenthis.writableEnded && this.writableLength === chunk.byteLength, the last queued chunk is processed with_finishFlushFlag, which ends the frame. ZlibBase#_flushthen calls_transformagain with an empty buffer.writableEndedis still true andwritableLengthis 0, so that call gets the finish flag too.
For deflate and brotli a second finish on a finished stream emits nothing. For zstd, ZSTD_compressStream2(..., ZSTD_e_end) on a context whose frame has just completed starts and completes a new frame. ZstdCompressContext::DoThreadPoolWork in src/node_zlib.cc doesn't guard against that. When end() is called with nothing queued, the frame is only finished once, so the output is a single frame.
The extra frame is valid zstd, but it has real consequences:
- The output depends on stream timing, not content. We hash the archive for deduplication, so identical inputs can hash differently.
- In released versions (including v26.9.0),
createZstdDecompressthrowsUnknown frame descriptorwhen a read chunk boundary falls inside those 9 bytes. Withfs.createReadStream's default 64 KiB chunks, that's about 1 archive in 8,200, and it fails the same way every time. We hit this in production on an archive thatzstd -tandzstdDecompressSyncboth accept. I believe #65865 fixes the decoding side onmain, but the compressor still writes the extra frame.
Our workaround is to strip a trailing 28b52ffd2000010000 after compressing.
- Ngôn ngữ chính
- JavaScript
- Star
- 122k
- Fork
- 37.4k
- Merge trung bình
- 4 ngày 4 giờ
- Pull request đã merge (30 ngày)
- 276
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của nodejs/node
-
doc
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
-
build
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
-
feature request
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Issue tương tự
-
bug customer-eng Durable Agents Inngest status: needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
optimization optimization:agents-md-curator
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
githubnext/gh-aw-cao#13475 ·
-
[BUG]: "Clear All" in Settings doesn't clear the saved analysis, old data comes back after reload Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
AOSSIE-Org/OrgExplorer#253 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
oxc-project/oxc#26944 ·
-
ai-observability bug team/ai-observability
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100