Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Older `parquet-java` readers fail on files with unprojected VARIANT columns

Open
#3,633 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
55/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Stale
Tech stack
java
Domain
data

Research direction

Start with ParquetMetadataConverter.getLogicalTypeAnnotation and the stack locations in ParquetMetadataConverter.java, then run the failing TestReadWithUnknownLogicalType test against parquet-testing/shredded_variant/case-001.parquet. Done means a projected read of id tolerates the unprojected VARIANT column while unsupported fields still fail only when semantically read.

Written by the indexing model from the issue text.

Description

Summary

parquet-java 1.15.x fails when opening a Parquet file that contains a VARIANT column, even if the application only requests other columns and does not request the VARIANT column.

The failure happens while converting the footer schema, before projection is applied. The parquet-format compatibility guidance says that new logical types are intended to be forward compatible: https://github.com/apache/parquet-format/blob/master/CONTRIBUTING.md#compatibility-and-feature-enablement

Based on that, I think older readers should be able to tolerate an unknown logical type when that field is not part of the requested projection.

Repro

This can be reproduced with the VARIANT fixture in apache/parquet-testing:

git clone https://github.com/apache/parquet-testing.git
PARQUET_FILE="$PWD/parquet-testing/shredded_variant/case-001.parquet"

The file contains an id column and a var column annotated with VARIANT.

Using parquet-java 1.15.x, try to read only id through the normal GroupReadSupport projection path:

Path input = new Path(args[0]);

try (ParquetReader<Group> reader = ParquetReader.builder(new GroupReadSupport(), input)
    .set(
        ReadSupport.PARQUET_READ_SCHEMA,
        "message root {\n" + "optional int32 id;\n" + "}")
    .build()) {
  Group row = reader.read();
  System.out.println(row.getInteger("id", 0));
}

I also put up a draft repro PR with a failing unit test against parquet-1.15.x: https://github.com/kevinjqliu/parquet-java/pull/1

That test uses apache/parquet-testing/shredded_variant/case-001.parquet and requests only the id column. It still fails while reading footer metadata, before the projected read can happen.

Actual behavior

The read fails during footer schema conversion:

[ERROR] org.apache.parquet.hadoop.TestReadWithUnknownLogicalType.testReadProjectedColumnFromFileWithUnknownLogicalType -- Time elapsed: 0.283 s <<< ERROR!
java.lang.NullPointerException: Cannot invoke "org.apache.parquet.format.LogicalType$_Fields.ordinal()" because the return value of "org.apache.parquet.format.LogicalType.getSetField()" is null
    at org.apache.parquet.format.converter.ParquetMetadataConverter.getLogicalTypeAnnotation(ParquetMetadataConverter.java:1174)
    at org.apache.parquet.format.converter.ParquetMetadataConverter.buildChildren(ParquetMetadataConverter.java:1892)
    at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetSchema(ParquetMetadataConverter.java:1840)
    at org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:1670)
    at org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:1630)
    at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:629)
    at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:934)
    at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:925)

The stack points to ParquetMetadataConverter.getLogicalTypeAnnotation.

Expected behavior

Older readers should tolerate an unknown logical type in the footer when the field is not projected, preserving the physical schema or treating the field as having no known logical annotation.

They should only fail if the unsupported field is actually read or interpreted semantically.

Notes

This is relevant for files written by newer or external writers that contain VARIANT, where older readers only need unrelated columns.

Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
4d 12h
Merged PRs (30d)
28

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/parquet-java

All issues in apache/parquet-java

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.