Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Isolate Avro usage in parquet-cli to commands that require Avro

オープン
#3,672 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
リファクタリング
明瞭さ
おおむね明確
活発さ
静か
技術スタック
java

調査の方向性

まず、既存の GroupReadSupport フォールバックを含め、cat、head、scan、schema、conversion、rewrite コマンドの parquet-cli エントリポイントを追跡します。data/nested_lists.snappy.parquet、shredded_variant/case-001.parquet、data/int96_from_spark.parquet で失敗を再現します。Parquet の検査と物理的な書き換えがネイティブパスを使用し、Avro は変換ワークフローと Parquet 以外の入力に引き続き使用されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Motivation

parquet-cli still uses Avro as an internal row model in places where users are only trying to inspect Parquet files. Valid Parquet is not always valid Avro, so Parquet inspection should not require Parquet schemas or records to round-trip through Avro.

cat/head already has a group-reader fallback for one class of problem: Avro schema conversion failures such as invalid Avro field names. That fallback is useful, but it is still a special case. Other inspection paths, such as scan, still use the Avro-backed data path for Parquet input.

Concrete examples:

  • Invalid Avro field names: original-instruction and column with known type are valid Parquet field names but invalid Avro names. This already required a cat/head group-reader fallback.
  • Nested Parquet structures: data/nested_lists.snappy.parquet is a valid parquet-testing fixture, but cat and scan fail through the Avro-backed reader.
  • Projection/type mismatch: UUID and INT96 Parquet columns can fail when Avro projection rewrites the requested physical type.
  • Lost Parquet structure: shredded Variant and other newer logical types can have valid Parquet physical structures that Avro cannot represent without losing information.

Avro is still part of the CLI API for conversion workflows. The goal is not to remove Avro; it is to stop requiring Avro for Parquet inspection.

Proposal

Use Avro only when the command or input/output format requires it. Use Parquet-native readers and metadata APIs for Parquet inspection and physical rewrites.

Command Avro needed? Proposed internal path
help, version No CLI metadata only
meta, pages, dictionary, check-stats, column-index, column-size, footer, bloom-filter, size-stats, geospatial-stats No ParquetFileReader and Parquet metadata/page/stat APIs
prune, trans-compression, masking, rewrite No Parquet rewrite utilities
cat, head Only for non-Parquet inputs Parquet -> GroupReadSupport; Avro/JSON -> Avro path
scan Only for non-Parquet inputs Parquet -> GroupReadSupport; Avro/JSON -> Avro path
schema Yes by default Keep Avro schema output; keep physical Parquet path for --parquet
csv-schema Yes Avro schema inference
convert-csv Yes CSV -> Avro records/schema -> Parquet
convert Yes Avro-backed conversion path
to-avro Yes Avro writer/schema path

Related issues

Tangential: https://github.com/apache/parquet-java/issues/2657 and https://github.com/apache/parquet-java/issues/2630 show additional CLI coupling to Avro internals, but are mostly binary/API mismatch issues.

Summary

Keep Avro for Avro-facing workflows: schema default output, to-avro, convert, CSV/JSON conversion, and Avro input handling.

Use Parquet-native readers for Parquet inspection and physical rewrites. This preserves the Avro API where intentional while making parquet-cli more robust for valid Parquet files that Avro cannot model.

Appendix
Repro examples

Using parquet-mr version 1.17.1 and fixtures from apache/parquet-testing:

git clone https://github.com/apache/parquet-testing.git
cd parquet-testing

These Parquet inspection commands currently fail:

parquet cat -n 1 data/nested_lists.snappy.parquet
parquet scan data/nested_lists.snappy.parquet

parquet cat -n 1 shredded_variant/case-001.parquet
parquet scan shredded_variant/case-001.parquet

parquet cat -n 1 data/int96_from_spark.parquet
parquet scan data/int96_from_spark.parquet

Observed failure classes:

  • data/nested_lists.snappy.parquet: Avro schema conversion succeeds, but cat and scan fail during Avro-backed record reading with ParquetDecodingException, followed by GroupColumnIO.getFirst index failure.
  • shredded_variant/case-001.parquet: Avro schema conversion succeeds, but cat and scan fail during Avro-backed record reading with ParquetDecodingException, followed by GroupColumnIO.getLast index failure.
  • data/int96_from_spark.parquet: Avro schema conversion fails before reading records with Argument error: INT96 is deprecated..., blocking Parquet row inspection.
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
4日 5時間
マージ済み PR(30日)
30

環境構築

  • Dockerfile・Docker Compose ファイルなし
  • プルリクエストのテンプレートあり
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/parquet-java のほかの issue

apache/parquet-java の issue をすべて見る

似ている issue

Java の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。