Decouple parquet-hadoop module from hadoop-mapreduce-client-core
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
- Tipo de issue
- Refactorización
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- hadoop, java
- Área
- data-engineering
Línea de trabajo
Comienza con ParquetReadOptions y ParquetInputFormat en el módulo parquet-hadoop y, a continuación, sigue las constantes existentes y la lógica de resolución de filtros. Introduce las clases org.apache.parquet.conf solicitadas, preservando los puntos de entrada existentes de ParquetInputFormat y la compatibilidad. Verifica que las nuevas clases y ParquetReadOptions no inicialicen FileInputFormat y, después, ejecuta ./mvnw test.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe the enhancement requested
Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).
Summary
To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.
ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.
Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.
ParquetInputFormat has three responsibilities:
- define a set of property keys as constants
- deserialize the filter predicates from a configuration value
- support the integration of Parquet files into Hadoop MapReduce
This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:
ParquetInputProperties: the property-key constants onlyParquetInputFilters: the filter deserialization logic
so that the configuration can be used without ever loadingFileInputFormatand its transitive dependencies.
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses
▼
ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)
▼
ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends
▼
FileInputFormat<Void, T> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends
▼
InputFormat<K, V> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)
▼
{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.
Public API / Behavioral change
No behavior change. This is a refactoring:
- New classes
org.apache.parquet.conf.ParquetInputPropertiesand
org.apache.parquet.conf.ParquetInputFilters(inparquet-hadoop) hold the constants and the
ParquetConfiguration-based filter resolution respectively. - The legacy
org.apache.hadoop.conf.Configuration-based entry points remain on
ParquetInputFormat(they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility. - All pre-existing constants on
ParquetInputFormatare now@Deprecatedand delegate to the new
classes; source and binary compatibility are preserved.
Acceptance criteria
ParquetReadOptions(used for plain file reads) references onlyParquetInputProperties/ParquetInputFilters, neverParquetInputFormat/org.apache.hadoop.mapreduce.InputFormat.- Loading
ParquetInputProperties/ParquetInputFilters/ParquetReadOptionsdoes not initializeFileInputFormat. - All existing tests still pass (
./mvnw test).
Component(s)
Core
- Lenguaje dominante
- Java
- Estrellas
- 3.1k
- Forks
- 1.6k
- Merge medio
- 4 d 5 h
- PR fusionados (30 d)
- 30
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/parquet-java
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
apache/parquet-java#3829 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
apache/parquet-java#3820 ·
Los mantenedores suelen responder en 2 días
-
Make PageReader AutoCloseableAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
apache/parquet-java#3767 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
apache/parquet-java#3695 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
apache/parquet-java#3667 ·
Los mantenedores suelen responder en 2 días
Todos los issues de apache/parquet-java
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
aoqia194/leaf-loader#19 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
apache/streampark#4521 ·
-
Update license yearAbierto0 - Backlog 1 - Ready documentation good first issue help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
-
cbor
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
FasterXML/jackson-dataformats-binary#844 ·
Los mantenedores suelen responder en 1 día
-
Issue: Bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
OpenAPITools/openapi-generator#25107 ·
Los mantenedores suelen responder en 1 día