Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

zlib: createZstdCompress appends an empty frame when end() is called with writes still queued

Offen
#66,078 0 Kommentare 1 Reaktion 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
68/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Bereich
backend, performance

Rechercherichtung

Beginne mit der bereitgestellten JavaScript-Reproduktion und lies anschließend lib/zlib.js, insbesondere ZlibBase#_transform und _flush. Untersuche ZstdCompressContext::DoThreadPoolWork in src/node_zlib.cc und vergleiche den Pfad für eingereihtes Schreiben mit dem Verhalten von gzip und brotli. Als abgeschlossen gilt die Aufgabe, wenn createZstdCompress bei einem oder mehreren Schreibvorgängen einen einzigen Frame mit identischen Bytes und keinen nachfolgenden leeren Frame ausgibt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Version

v26.9.0 (also v24.15.0)

Platform
Darwin 25.6.0 arm64 (also seen on Linux x64)
Subsystem

zlib

What steps will reproduce the bug?
'use strict';
const zlib = require('node:zlib');

function compress(create, writes) {
  return new Promise((resolve, reject) => {
    const stream = create();
    const chunks = [];
    stream.on('data', (chunk) => chunks.push(chunk));
    stream.on('end', () => resolve(Buffer.concat(chunks)));
    stream.on('error', reject);
    for (const w of writes) stream.write(w);
    stream.end();
  });
}

(async () => {
  const oneWrite = await compress(zlib.createZstdCompress, ['hello world']);
  const twoWrites = await compress(zlib.createZstdCompress, ['hello ', 'world']);

  console.log('one write: ', oneWrite.toString('hex'));
  console.log('two writes:', twoWrites.toString('hex'));
  console.log('extra bytes:', twoWrites.subarray(oneWrite.length).toString('hex'));

  // Same input, queued writes: gzip and brotli are unaffected.
  for (const [name, create, decompress] of [
    ['gzip', zlib.createGzip, zlib.gunzipSync],
    ['brotli', zlib.createBrotliCompress, zlib.brotliDecompressSync],
  ]) {
    const a = await compress(create, ['hello world']);
    const b = await compress(create, ['hello ', 'world']);
    console.log(name, 'lengths', a.length, b.length, decompress(b).toString());
  }
})();
How often does it reproduce? Is there a required condition?

Every time end() is called while at least one write is still queued in the compressor. In practice that is most pipelines whose source ends quickly, e.g. pipeline(tar.create(...), zlib.createZstdCompress(), fs.createWriteStream(...)): every archive we packed that way had the extra frame.

What is the expected behavior? Why is that the expected behavior?

One zstd frame, the same bytes whether the input arrived in one write or several:

one write:  28b52ffd005859000068656c6c6f20776f726c64
two writes: 28b52ffd005859000068656c6c6f20776f726c64

That's what gzip and brotli do. How the input was split into writes shouldn't change the compressed output.

What do you see instead?
one write:  28b52ffd005859000068656c6c6f20776f726c64
two writes: 28b52ffd005859000068656c6c6f20776f726c6428b52ffd2000010000
extra bytes: 28b52ffd2000010000
gzip lengths 31 31 hello world
brotli lengths 15 15 hello world

A second, empty zstd frame (28b52ffd 20 00 01 00 00: magic, single-segment descriptor with content size 0, one empty raw last block) is appended.

Additional information

I think the cause is in lib/zlib.js:

  • In ZlibBase#_transform, when this.writableEnded && this.writableLength === chunk.byteLength, the last queued chunk is processed with _finishFlushFlag, which ends the frame.
  • ZlibBase#_flush then calls _transform again with an empty buffer. writableEnded is still true and writableLength is 0, so that call gets the finish flag too.

For deflate and brotli a second finish on a finished stream emits nothing. For zstd, ZSTD_compressStream2(..., ZSTD_e_end) on a context whose frame has just completed starts and completes a new frame. ZstdCompressContext::DoThreadPoolWork in src/node_zlib.cc doesn't guard against that. When end() is called with nothing queued, the frame is only finished once, so the output is a single frame.

The extra frame is valid zstd, but it has real consequences:

  1. The output depends on stream timing, not content. We hash the archive for deduplication, so identical inputs can hash differently.
  2. In released versions (including v26.9.0), createZstdDecompress throws Unknown frame descriptor when a read chunk boundary falls inside those 9 bytes. With fs.createReadStream's default 64 KiB chunks, that's about 1 archive in 8,200, and it fails the same way every time. We hit this in production on an archive that zstd -t and zstdDecompressSync both accept. I believe #65865 fixes the decoding side on main, but the compressor still writes the extra frame.

Our workaround is to strip a trailing 28b52ffd2000010000 after compressing.

Vorherrschende Sprache
JavaScript
Sterne
122k
Forks
37.4k
Ø Merge
4 T. 12 Std.
Gemergte PRs (30 T.)
296

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus nodejs/node

Alle Issues in nodejs/node

Ähnliche Issues

Weitere Issues zu JavaScript

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.