Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)"

Open Beginner friendly
#6,251 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
scala, spark
Domain
backend

Research direction

Start at CometArraysZip.getSupportLevel and compare its checks with CometCreateNamedStruct.getSupportLevel. Reproduce the three arrays_zip queries from the issue, then make duplicate field names fall back rather than reach JVM Arrow import. Confirm arrays_zip(a, b) still runs natively and the two same-named cases match Spark without the import error.

Written by the indexing model from the issue text.

Description

bug requires-triage
Describe the bug

arrays_zip over two inputs with the same name fails the task when Comet runs it natively. Spark names each struct field after its input, so arrays_zip(a, a) has type array<struct<a: string, a: string>>, and Comet's native projection returns that struct fine. Importing the result on the JVM then fails, because Java Arrow keys struct children by name and the two a children collapse into one:

java.lang.IllegalStateException: ArrowArray struct has 2 children (expected 1)
	at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
	at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
	at org.apache.arrow.c.ArrayImporter.importChild(ArrayImporter.java:83)
	at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:101)
	at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
	at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
	at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
	at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
	at org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:237)

arrays_zip(a, b) over the same table works. Spark returns the zipped rows for all three queries below.

This is the same Java Arrow limitation as #1015 and #5605. CometCreateNamedStruct.getSupportLevel already returns Unsupported when names has duplicates, and DataTypeSupport rejects such structs in operator schemas, but CometArraysZip.getSupportLevel only checks the input types and never looks at expr.names. Two same-named inputs come up naturally after a join, for example arrays_zip(t1.tags, t2.tags).

Steps to reproduce

Default configs, reproduced on main at f7952de73 with Spark 4.1.3:

spark.range(4)
  .selectExpr("id", "array(cast(id as string), 'x') as a", "array(id, id + 1) as b")
  .write.parquet(path)
val df = spark.read.parquet(path)
df.selectExpr("id", "arrays_zip(a, a) AS r").collect() // IllegalStateException
df.selectExpr("id", "arrays_zip(b, b) AS r").collect() // IllegalStateException
df.selectExpr("id", "arrays_zip(a, b) AS r").collect() // matches Spark

The plan is CometProject over CometNativeScan.

Expected behavior

Either match Spark or fall back. The smallest fix is probably for CometArraysZip.getSupportLevel to return Unsupported when expr.names has duplicates, the same way CometCreateNamedStruct does.

Additional context

Found while reviewing #6036, but unrelated to that PR's change.

Dominant language
Scala
Stars
1.3k
Forks
377
Avg merge
2d 8h
Merged PRs (30d)
271

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/datafusion-comet

All issues in apache/datafusion-comet

Similar issues

More Scala issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.