JSONL File Scan reads a JSON null as the text "null"
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 86/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- scala
- Domain
- data-engineering
Research direction
Start at JSONUtils.JSONToMap and inspect the value-node handling for JSON null. Run JSONUtilsSpec, including its two assertions for null, and verify the JSONL File Scan reproduction with {"a":1} and {"a":null}. Done means null remains an empty value and schema inference produces an INTEGER column rather than STRING.
Written by the indexing model from the issue text.
Description
What happened?
The JSONL File Scan reads a JSON null as the text "null". One null in a column is enough to make the whole column STRING, so a column of numbers comes out as text and its missing value is the word null rather than an empty cell. A sort, a filter or any arithmetic downstream then works on the wrong type.
The cause is JSONUtils.JSONToMap: a null is a value node, so it takes the child.asText() branch, and NullNode.asText() returns "null". Schema inference then fails to parse "null" as a number and widens the column to STRING.
Two assertions in JSONUtilsSpec expect "null" today. They were added in #4716, whose description says the spec pins the current behavior so that a fix has to update it on purpose.
How to reproduce?
- Upload a JSONL file with these two lines:
{"a":1} {"a":null} - Add a JSONL File Scan that reads it and run the workflow.
- Column
ais inferred as STRING, and the two rows hold the text"1"and the text"null". Expected: INTEGER, with1and an empty value.
Version/Branch
1.4.0-incubating-SNAPSHOT (main)
Commit Hash (Optional)
53052c482
What browsers are you seeing the problem on?
No response
Relevant log output
Schema[Attribute[name=a, type=string]]
row 1: '1' (String)
row 2: 'null' (String)
- Dominant language
- Scala
- Stars
- 316
- Forks
- 189
- Avg merge
- 4d 21h
- Merged PRs (30d)
- 162
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/texera
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/texera#8700 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
apache/texera#8682 · 4 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
apache/texera#8300 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
apache/texera#8235 · 2 comments ·
Maintainers usually reply within 1 day
Similar issues
-
bug performance priority:medium regression
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/datafusion-comet#6313 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
com-lihaoyi/mill#7624 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
lichess-org/lila#21845 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day