Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Decouple parquet-hadoop module from hadoop-mapreduce-client-core

Open
#3,780 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
68/100
Issue type
Refactor
Clarity
Clearly specified
Activity status
Active
Tech stack
hadoop, java

Research direction

Start with ParquetReadOptions and ParquetInputFormat in the parquet-hadoop module, then trace the existing constants and filter-resolution logic. Introduce the requested org.apache.parquet.conf classes while preserving the legacy ParquetInputFormat entry points and compatibility. Verify that the new classes and ParquetReadOptions do not initialize FileInputFormat, then run ./mvnw test.

Written by the indexing model from the issue text.

Description

Type: enhancement
Describe the enhancement requested

Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).

Summary

To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.

ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.

Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.

ParquetInputFormat has three responsibilities:

  • define a set of property keys as constants
  • deserialize the filter predicates from a configuration value
  • support the integration of Parquet files into Hadoop MapReduce

This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:

  • ParquetInputProperties: the property-key constants only
  • ParquetInputFilters: the filter deserialization logic
    so that the configuration can be used without ever loading FileInputFormat and its transitive dependencies.
ParquetReader                                   ← org.apache.parquet.hadoop (parquet-hadoop)
        │  uses
        ▼
ParquetReadOptions                              ← org.apache.parquet (parquet-hadoop)
        │  (static import keys + getFilter call)
        ▼
ParquetInputFormat.getFilter(...)               ← org.apache.parquet.hadoop (parquet-hadoop)
        │  extends
        ▼
FileInputFormat<Void, T>                        ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  extends
        ▼
InputFormat<K, V>                               ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  (transitive)
        ▼
{ InputSplit, JobContext, TaskAttemptContext,
  RecordReader, ... }                           ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)

With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.

Public API / Behavioral change

No behavior change. This is a refactoring:

  • New classes org.apache.parquet.conf.ParquetInputProperties and
    org.apache.parquet.conf.ParquetInputFilters (in parquet-hadoop) hold the constants and the
    ParquetConfiguration-based filter resolution respectively.
  • The legacy org.apache.hadoop.conf.Configuration-based entry points remain on
    ParquetInputFormat (they are used only via the legacy MapReduce path) and are kept for binary /
    source compatibility.
  • All pre-existing constants on ParquetInputFormat are now @Deprecated and delegate to the new
    classes; source and binary compatibility are preserved.

Acceptance criteria

  • ParquetReadOptions (used for plain file reads) references only ParquetInputProperties / ParquetInputFilters, never ParquetInputFormat / org.apache.hadoop.mapreduce.InputFormat.
  • Loading ParquetInputProperties / ParquetInputFilters / ParquetReadOptions does not initialize FileInputFormat.
  • All existing tests still pass (./mvnw test).
Component(s)

Core

Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
4d 12h
Merged PRs (30d)
28

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/parquet-java

All issues in apache/parquet-java

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.