Decouple parquet-hadoop module from hadoop-mapreduce-client-core
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
- Issue type
- Refactor
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- hadoop, java
- Domain
- data-engineering
Research direction
Start with ParquetReadOptions and ParquetInputFormat in the parquet-hadoop module, then trace the existing constants and filter-resolution logic. Introduce the requested org.apache.parquet.conf classes while preserving the legacy ParquetInputFormat entry points and compatibility. Verify that the new classes and ParquetReadOptions do not initialize FileInputFormat, then run ./mvnw test.
Written by the indexing model from the issue text.
Description
Describe the enhancement requested
Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).
Summary
To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.
ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.
Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.
ParquetInputFormat has three responsibilities:
- define a set of property keys as constants
- deserialize the filter predicates from a configuration value
- support the integration of Parquet files into Hadoop MapReduce
This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:
ParquetInputProperties: the property-key constants onlyParquetInputFilters: the filter deserialization logic
so that the configuration can be used without ever loadingFileInputFormatand its transitive dependencies.
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses
▼
ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)
▼
ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends
▼
FileInputFormat<Void, T> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends
▼
InputFormat<K, V> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)
▼
{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.
Public API / Behavioral change
No behavior change. This is a refactoring:
- New classes
org.apache.parquet.conf.ParquetInputPropertiesand
org.apache.parquet.conf.ParquetInputFilters(inparquet-hadoop) hold the constants and the
ParquetConfiguration-based filter resolution respectively. - The legacy
org.apache.hadoop.conf.Configuration-based entry points remain on
ParquetInputFormat(they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility. - All pre-existing constants on
ParquetInputFormatare now@Deprecatedand delegate to the new
classes; source and binary compatibility are preserved.
Acceptance criteria
ParquetReadOptions(used for plain file reads) references onlyParquetInputProperties/ParquetInputFilters, neverParquetInputFormat/org.apache.hadoop.mapreduce.InputFormat.- Loading
ParquetInputProperties/ParquetInputFilters/ParquetReadOptionsdoes not initializeFileInputFormat. - All existing tests still pass (
./mvnw test).
Component(s)
Core
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 4d 12h
- Merged PRs (30d)
- 28
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/parquet-java
-
Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
apache/parquet-java#3792 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
apache/parquet-java#3767 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
apache/parquet-java#3695 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/parquet-java#3667 ·
-
Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
apache/parquet-java#3587 ·
All issues in apache/parquet-java
Similar issues
-
awaiting triage bug Causes friction Hop Gui P1 P2 Transforms
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
apache/flink-agents#1152 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
jenkinsci/blueocean-plugin#5417 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
objectionary/eo-graphs#75 ·