[Feature] Prefetch the late-materialization payload ranges instead of reading them on demand
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- cpp
- Área
- data-engineering, performance
Línea de trabajo
Comienza siguiendo PrefetchFileBatchReader::PreBufferRange(), ReadAheadCache::Init(), Read(), Reset() y Close(), y después inspecciona las APIs públicas bajo include/paimon/. Sigue cómo LateMaterializingFileBatchReader determina los rangos de payload y cómo se registran las métricas de caché existentes. La tarea estará terminada cuando los rangos tardíos se registren y se calienten de forma segura entre rondas, se descarten los registros obsoletos, se cumplan las reglas de concurrencia y solapamiento, y las nuevas métricas contabilicen los bytes registrados y descartados.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Late materialization reads a data file in two passes: a probe pass over the predicate fields, then a payload pass over the remaining fields for the matched rows only. The shared read-ahead cache is fed once per read-range generation through PrefetchFileBatchReader::PreBufferRange(), before any read starts. At that point the payload pass cannot know which pages hold the matched rows — that depends on the probe result — so PreBufferRange() only reports the probe ranges (there was an explicit TODO for exactly this). The payload pass is therefore never prefetched: every payload read misses the cache and waits for its own underlying IO, serialized against the decode, on the pass that touches the wide columns.
Solution
Let a reader report byte ranges that only become known after reading has started, and let the shared cache register them mid-read.
- New
ReadAheadCache::AddRanges(ranges, expected_round)registers ranges into an already-initialized cache and is safe to call repeatedly and concurrently withRead(). It merges the new ranges into the disjoint, offset-ordered pending list, registering only the parts no registered range covers and dropping the overlap (the round that registered it is already fetching those bytes), then rebuilds the per-range cached flags so an already-fetched range is not fetched twice. The registered part is cut at a newCacheConfigknoblate_range_size_limit(default 8 MiB, smaller than the 32 MiBrange_size_limit) so a large pass is fetched by several concurrent requests rather than one long one; a newWarmup(from_offset)starts fetching from the first newly-registered range instead of from the head. - A registration round bounds the lifetime: every
Init()opens a round identified byRegistrationRound(), andAddRanges()drops everything whenexpected_roundis not the open round, so a pass that outlived its generation — the cache was reset for a new read-range generation, or released byClose()— registers nothing instead of prefetching bytes nobody reads. The round counter is monotonic acrossReset()so a stale round is never mistaken for a new one. - New
PrefetchFileBatchReader::PreBufferSinkandSetPreBufferSink():PrefetchFileBatchReaderImplinstalls a sink on each sub-reader that tags the reported ranges with the current round, callsAddRanges, and warms up from the first new range.LateMaterializingFileBatchReaderreports the payload ranges through the sink once the probe pass has refined the inner reader's target pages, and surfaces a failure to compute them (they come from the file metadata) rather than swallowing it. - New metrics
read-ahead-cache.late.registered/.registered-bytes/.dropped/.dropped-bytes, counted after coalescing and splitting, soregistered-bytesanddropped-bytestogether account for every reported byte.
Anything else?
Adds public API under include/paimon/: PrefetchFileBatchReader::PreBufferSink / SetPreBufferSink() and CacheConfig::GetLateRangeSizeLimit() / SetLateRangeSizeLimit(). ReadAheadCache::AddRanges / RegistrationRound / Warmup(offset) and the new counter names live in the internal header. No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Lenguaje dominante
- C++
- Estrellas
- 65
- Forks
- 29
- Merge medio
- 2 d 4 h
- PR fusionados (30 d)
- 78
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/paimon-cpp
-
enhancement
apache/paimon-cpp#381 · 1 asignado ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 30/100
apache/paimon-cpp#375 · 1 asignado ·
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 45/100
apache/paimon-cpp#361 · 1 asignado ·
-
enhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
apache/paimon-cpp#325 · 1 asignado ·
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
apache/paimon-cpp#319 · 1 reacción · 1 asignado ·
Todos los issues de apache/paimon-cpp
Issues similares
-
AuTest Bug Tests
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
apache/trafficserver#13714 ·
-
bug build
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
facebookincubator/velox#19143 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
tenstorrent/tt-metal#57393 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
objectionary/eo-graphs#74 ·