Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Vectored Parquet reads can fall back unsafely after partial asynchronous reads and exceed allocation limits

Aperta
#3,719 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
45/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Tranquilla
Stack tecnologico
java

Direzione di ricerca

Inizia in parquet-hadoop's ParquetFileReader.java nel percorso di lettura vettoriale intorno alle righe 1293-1307, quindi segui ChunkListBuilder, le letture asincrone dei fratelli e la gestione di parquet.read.allocation.size. Aggiungi una copertura di regressione per l'invio o il completamento parziali, i futures dei fratelli in sospeso, le pagine filtrate, le colonne sovradimensionate e le letture con checksum abilitato; il lavoro è completato quando sono garantiti un fallimento sicuro o un fallback e allocazioni limitate, senza modificare i risultati decodificati.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Describe the bug, including details regarding any error messages, version, and platform.

ParquetFileReader can produce unsafe fallback behavior when Hadoop vectored I/O
is enabled and a filesystem partially submits or completes a vectored read before
raising IllegalArgumentException or UnsupportedOperationException.

The current implementation catches those exceptions around readVectored(...)
and retries every range using ordinary reads against the same ChunkListBuilder.
If an earlier range already populated the builder, its data is appended again.
For a filtered column with selected pages P0 and P2, the buffered page
sequence can become [P0, P0, P2] although the page index still describes
[P0, P2]. Depending on the page contents, decoding can fail or silently
associate the wrong page with the selected rows. Even when no page has been
consumed yet, scalar fallback is unsafe once sibling asynchronous reads may still
be operating on the same stream.

Current upstream code:

https://github.com/apache/parquet-java/blob/8e30c4cee3c7e85a8cf2133697f13138509b05b7/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java#L1293-L1307

The same vectored path also allocates one buffer for an entire contiguous
requested range instead of honoring parquet.read.allocation.size (8 MiB by
default), unlike the ordinary read path. Large projected column chunks or
filtered pages can therefore create unexpectedly large heap allocations.

The current master branch and Apache Parquet Java 1.18.0 contain this behavior.
The vulnerable path is also present in the 1.15.x, 1.16.x, and 1.17.x release
lines. Vectored I/O defaults to enabled starting in 1.16.0, so supported Hadoop
filesystems can reach this path without an explicit opt-in; in 1.15.2 it is
reachable when explicitly enabled.

Expected behavior:

  • Preserve ordinary fallback only when vectored I/O is unavailable or range
    preparation fails before asynchronous submission starts.
  • Once submission begins, fail the read safely rather than replaying scalar reads
    against a partially populated builder or an active stream.
  • Wait for already-published sibling reads before returning the original failure.
  • Split filesystem byte ranges to respect the configured allocation limit without
    changing the logical read plan or decoded results.
  • Add regression coverage for partial submission/completion, pending sibling
    futures, filtered pages, oversized columns, and checksum-enabled reads.
Component(s)

parquet-hadoop

Lingua principale
Java
Stelle
3.1k
Fork
1.6k
Merge medio
6g 16h
PR unite (30g)
36

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/parquet-java

Tutte le issue di apache/parquet-java

Issue simili

Altre issue su Java

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.