Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Feature] Warm up next data file in ConcatBatchReader

Aperta
#419 0 commenti 0 reazioni 1 assegnatario Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@SteNicholas ci sta già lavorando.

Dal 4/10/2026.

  • #420 di @SteNicholas — aperta

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

enhancement

Search before asking

  • I searched in the issues and found nothing similar.

Motivation

#289 overlapped the first read of a data file with the consumption of the previous one, but only on the merge-on-read path: it added Warmup() to FileBatchReader and KeyValueRecordReader and calls it from ConcatKeyValueRecordReader and LoserTree. The read paths that do not merge concatenate their data files through ConcatBatchReader instead, which never calls Warmup(), so they still issue the first remote read of every file serially, only after the previous file has reached EOF.

The most direct case is append compaction. AppendOnlyFileStoreWrite::CreateFilesReader builds its read context with prefetch enabled and leaves the read-ahead cache and the warmup level at their defaults (enabled and WarmupLevel::RAW), then reads the files to compact through RawFileSplitRead::CreateReader, which is one ConcatBatchReader over the files' FileBatchReader stacks. CompactRewrite drains that reader sequentially on a single thread. Every file's reader is already built and has its read schema set before the first batch is read, because CreateRawFileReadersWithMeta builds the readers of all files up front, so every file that is actually prefetched could be warmed as is; nothing ever asks. A compaction that rewrites many small files, which is the rewrite append compaction exists for, therefore waits for one remote round trip at every such file boundary. Not every compacted file is prefetched, though; see the adaptive-strategy limit below.

The same ConcatBatchReader concatenates FileBatchReaders on other paths too, which would get the same overlap:

  • append-table batch reads, through the same RawFileSplitRead::CreateReader that compaction uses;
  • raw-convertible primary-key splits read without merging, through the raw path of MergeFileSplitRead;
  • the blob files of one data file in DataEvolutionSplitRead, where it is a no-op in practice because blob does not go through the prefetch reader.

Solution

Warm one file ahead in ConcatBatchReader, the same way ConcatKeyValueRecordReader does since #289.

  • ConcatBatchReader holds its children as BatchReader, while Warmup() is declared on FileBatchReader. Resolve each child's FileBatchReader* once in the constructor, nullptr for a child that is not one, which is the same boundary cast KeyValueDataFileRecordReader::Warmup() makes, rather than casting on every batch. Children that are not FileBatchReaders keep today's behavior: the per-split readers concatenated by AppendOnlyTableRead and KeyValueTableRead, the fallback and audit-log readers, the sort-merge sections of MergeFileSplitRead, and the realtime commit readers.
  • In NextBatchWithBitmap(), warm the current child and kWarmupLookahead = 1 child after it before reading, as ConcatKeyValueRecordReader::NextBatch() does. Warmup() is idempotent, so on readers that are already warm this costs one virtual call and one pointer test per child per batch. NextBatch() goes through NextBatchWithBitmap() and needs no change of its own.
  • Do not add Warmup() to the public BatchReader. That would change the vtable of an exported class, and it would only be needed to warm a ConcatBatchReader from the outside, which no caller does today.

Tests go into the existing concat_batch_reader_test.cpp, with a MockFileBatchReader that logs its Warmup() and read calls, in order, into a log the test owns, since a child is destroyed once it reaches EOF: nothing is warmed before the first read, child i + 1 is warmed before child i starts reading and nothing beyond it is, the lookahead window counts every child so a skipped child does not widen it, a child that is not a FileBatchReader is never warmed, and the output is unchanged.

Document warmup in the prefetch user guide (docs/source/user_guide/prefetch.rst): the levels, where it applies, the adaptive-strategy and synchronous-file-system limits below, and the memory cost.

Anything else?

Warming changes when a file's first bytes are fetched, never what is read: ordering, filtering, deletion vectors and metrics are untouched. As in #289 there is no new configuration; the existing ReadContextBuilder::SetWarmupLevel() and SetReadAheadCacheEnabled() already control it. It only has an effect for formats that go through the prefetch reader, so avro, blob, lance and mosaic files are unaffected.

The costs and limits, stated up front:

  • Memory. Under RAW the next file's ranges are fetched through its read-ahead cache, up to CacheConfig pre_buffer_limit (256 MiB by default), so while one file is being consumed the next one can hold up to min(its data size, 256 MiB) on top of it. Append compaction builds its own ReadContext, and the warmup level is not a table option, so compaction cannot opt out. This is the trade-off #289 and #365 already accepted for merge reads; if it turns out to matter for compaction, CreateFilesReader can set the level explicitly in a follow-up rather than adding a knob here.
  • Adaptive strategy. Warmup only reaches files that are actually prefetched, and enabling prefetch does not guarantee that. With the adaptive strategy, which is on by default, PrefetchFileBatchReaderImpl reads a file without prefetching when its first read range holds more batches than one prefetch queue can buffer, and DelegatingPrefetchReader::Warmup() then returns without doing anything. Append compaction reads with one prefetch reader and a queue of three batches, and a Parquet read range is a whole row group, so with the default 1024-row batches any row group of 4096 rows or more is read without prefetching and its file is not warmed. Most append compactions on Parquet therefore gain little from this change; files that do go through prefetch, typically ones with small row groups, still do. Making compaction benefit more, by giving it a larger prefetch queue or by letting RAW warm a reader the adaptive strategy bypassed, is a separate change with its own memory trade-off and is left as a follow-up.
  • Synchronous file systems. RAW warmup issues its fetches from the reading thread through InputStream::ReadAsync, so it only overlaps I/O when that call is really asynchronous, as it is for Jindo and the object stores. LocalInputStream::ReadAsync runs the read synchronously, so on the local file system RAW warmup reads up to the pre-buffer limit of the next file on the reading thread before the current batch is returned, adding latency instead of hiding it. DECODED fetches on the background decode thread and is not affected. The merge-on-read path has behaved this way since #289; this change extends it to the non-merging paths. Avoiding it needs either a way to tell synchronous streams apart or RAW warmup dispatched off the reading thread, as Velox does by preparing preloaded splits on an IO executor; either changes more than ConcatBatchReader and is left as a follow-up. The user guide documents the limit and points to DECODED or NONE for such reads.
  • Gain. Warming saves the first-read round trip at each boundary into a prefetched file. It matters for compactions and splits made of many small files, and is marginal for a few large ones.
  • Early stop. A read that stops early, for example under a LIMIT, may already have warmed a file it never reads. ReadAheadCache::ReleasePrefetchBuffers() waits for dispatched fetches before freeing their buffers, so closing such a read, or cancelling a compaction, can also wait for the warmed file's in-flight fetches. The merge path has the same behavior today.

Warming still does not cross a split boundary, because the outer ConcatBatchReader of AppendOnlyTableRead and KeyValueTableRead concatenates per-split readers that are not FileBatchReaders. That is left as a follow-up. For comparison, Velox preloads whole splits ahead of the consumer (max_split_preload_per_driver, 2 by default) and opens the file and its reader on an IO executor. Paimon already builds every reader of a split up front, so within a split only the first data fetch remains to be overlapped, and that is what this change covers.

Are you willing to submit a PR?

  • I'm willing to submit a PR!
Lingua principale
C++
Stelle
65
Fork
31
Merge medio
1g 14h
PR unite (30g)
60

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/paimon-cpp

Tutte le issue di apache/paimon-cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.