Expose Parquet Data Page dictionary index provenance to downstream readers
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- rust
- Ambito
- data-engineering, databases
Direzione di ricerca
Start by reviewing ColumnReader::read_records, DictIndexDecoder, and the private parquet::arrow::decoder and experimental encodings::rle paths. Compare the proposed callback and value-section decoder against the V1/V2 framing, null-count, fallback, truncation, and dictionary-ID requirements, including the fix from #10725. Done means maintainers agree on a supported public seam with a focused implementation and validation scope.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
A physical Parquet Page inspector needs the encoded dictionary entry ID consumed by each non-null value position of one RLE_DICTIONARY Data Page. This is different from the occurrence ordinal and from returning an Arrow DictionaryArray: decoded values can repeat while their occurrence ordinals advance, for example ordinals 0, 1, 2 consuming dictionary IDs 2, 2, 0.
ColumnReader::read_records returns repetition/definition levels and already materialized values, but does not expose the consumed dictionary IDs. DictIndexDecoder would decode the values section, but parquet::arrow::decoder is private (including on current main). The lower-level encodings::rle route is available only via the experimental feature, which carries no stability guarantee. This was verified against parquet 58.4.0 with downstream imports failing E0603. We are looking for a supported seam rather than making all decoder internals public.
Related #9010 asks to inspect dictionary contents in the async reader; this request is for per-occurrence IDs in a selected Data Page.
Possible interface
Would maintainers prefer either:
- an opt-in Page/column decode callback that pairs each level position with the optional dictionary ID and decoded value; or
- a public, validated value-section dictionary-index decoder with an explicit count contract, documented for composition with levels from
ColumnReader?
The second route would let an inspector split the already decompressed selected Page into V1/V2 level and value sections, then align returned IDs only to positions whose definition level equals the schema maximum. A Page can fall back to PLAIN; those positions must have no dictionary ID. The API should permit bounded, opt-in use without adding per-value work to ordinary reads.
Correctness boundary
- V1 and V2 level-section framing differ; a caller must isolate the values section before decoding indices.
- The number of IDs must equal the number of non-null physical values, including for nested repetition/definition levels and V2
num_nulls. - Reject empty/truncated streams, invalid bit widths, and negative or out-of-range dictionary IDs. Current
mainhas the empty-section/bit-width fix from #10725; an exposed seam should retain it. - No dictionary ID can be inferred by searching a decoded value in the dictionary, nor from its occurrence ordinal.
- A
PLAINfallback Page must report no dictionary ID even when an earlier Page in the Column Chunk used a dictionary.
Could you advise which public interface fits parquet-rs ownership and stability expectations? We can contribute a focused implementation after the seam is agreed.
- Lingua principale
- Rust
- Stelle
- 3.6k
- Fork
- 1.3k
- Merge medio
- 3g 10h
- PR unite (30g)
- 148
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/arrow-rs
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
apache/arrow-rs#11225 · 2 commenti · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
I maintainer di solito rispondono entro 1 giorno
-
parquet-variant-compute: shred_variant panics on an object with duplicate field namesForse già presa @Abhisheklearn12 l’ha presa 14 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di apache/arrow-rs
Issue simili
-
bug CLI exec tool-calls
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
maintainer-needed p2 triaged ui windows
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
ai_p2
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
ClickHouse/ClickHouse#123351 ·
I maintainer di solito rispondono entro 1 giorno
-
documentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
github/copilot-sdk#2804 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno