Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Decouple parquet-hadoop module from hadoop-mapreduce-client-core

Abierto
#3,780 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 2 días

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
68/100
Tipo de issue
Refactorización
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
hadoop, java

Línea de trabajo

Comienza con ParquetReadOptions y ParquetInputFormat en el módulo parquet-hadoop y, a continuación, sigue las constantes existentes y la lógica de resolución de filtros. Introduce las clases org.apache.parquet.conf solicitadas, preservando los puntos de entrada existentes de ParquetInputFormat y la compatibilidad. Verifica que las nuevas clases y ParquetReadOptions no inicialicen FileInputFormat y, después, ejecuta ./mvnw test.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Type: enhancement
Describe the enhancement requested

Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).

Summary

To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.

ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.

Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.

ParquetInputFormat has three responsibilities:

  • define a set of property keys as constants
  • deserialize the filter predicates from a configuration value
  • support the integration of Parquet files into Hadoop MapReduce

This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:

  • ParquetInputProperties: the property-key constants only
  • ParquetInputFilters: the filter deserialization logic
    so that the configuration can be used without ever loading FileInputFormat and its transitive dependencies.
ParquetReader                                   ← org.apache.parquet.hadoop (parquet-hadoop)
        │  uses
        ▼
ParquetReadOptions                              ← org.apache.parquet (parquet-hadoop)
        │  (static import keys + getFilter call)
        ▼
ParquetInputFormat.getFilter(...)               ← org.apache.parquet.hadoop (parquet-hadoop)
        │  extends
        ▼
FileInputFormat<Void, T>                        ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  extends
        ▼
InputFormat<K, V>                               ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  (transitive)
        ▼
{ InputSplit, JobContext, TaskAttemptContext,
  RecordReader, ... }                           ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)

With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.

Public API / Behavioral change

No behavior change. This is a refactoring:

  • New classes org.apache.parquet.conf.ParquetInputProperties and
    org.apache.parquet.conf.ParquetInputFilters (in parquet-hadoop) hold the constants and the
    ParquetConfiguration-based filter resolution respectively.
  • The legacy org.apache.hadoop.conf.Configuration-based entry points remain on
    ParquetInputFormat (they are used only via the legacy MapReduce path) and are kept for binary /
    source compatibility.
  • All pre-existing constants on ParquetInputFormat are now @Deprecated and delegate to the new
    classes; source and binary compatibility are preserved.

Acceptance criteria

  • ParquetReadOptions (used for plain file reads) references only ParquetInputProperties / ParquetInputFilters, never ParquetInputFormat / org.apache.hadoop.mapreduce.InputFormat.
  • Loading ParquetInputProperties / ParquetInputFilters / ParquetReadOptions does not initialize FileInputFormat.
  • All existing tests still pass (./mvnw test).
Component(s)

Core

Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.6k
Merge medio
4 d 5 h
PR fusionados (30 d)
30

Preparar el entorno

  • Sin Dockerfile ni archivo de Docker Compose
  • Tiene una plantilla de pull request
  • Sin guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de apache/parquet-java

Todos los issues de apache/parquet-java

Issues similares

Más issues de Java

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.