CsvReadOptions.null_regex has no effect: matching values are read as literal strings, and fail the read in a numeric column
還沒有人認領這個 Issue。
評估
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 新手友好度
- 76/100
- Issue 類型
- 缺陷
- 描述清晰度
- 描述清楚
- 活躍度
- 活躍
研究方向
從 crates/core/src/options.rs 和 datafusion-datasource-csv 的 src/source.rs:187 開始,然後將 reader builder 與 src/file_format.rs:547 中對 null_regex 的處理進行比較。在 python/tests/test_context.py 中使用數值欄中的相符值新增覆蓋,並執行 CSV 選項測試。當相符欄位在字串欄和數值欄中都變為 NULL,且 docs/source/user-guide/io/csv.md 中記載的行為仍然正確時,即可完成。
由索引模型根據 Issue 內容生成。
描述
Describe the bug
CsvReadOptions.null_regex is accepted everywhere it is offered, but the CSV
reader never applies it. Values matching the regex are read as literal strings,
and when a matching value sits in a column typed as an integer the read fails
outright instead of producing NULL.
This is not a bug in this crate — python/datafusion/options.py stores the
value and crates/core/src/options.rs copies it into DataFusion's
CsvReadOptions correctly. The root cause is in DataFusion itself; see below.
To Reproduce
from pathlib import Path
from datafusion import CsvReadOptions, SessionContext
p = Path("probe.csv")
p.write_text("id,name\n1,alice\n2,N/A\n3,carol\n")
ctx = SessionContext()
options = CsvReadOptions().with_has_header(True).with_null_regex(r"^(null|NULL|N/A)$")
ctx.read_csv(p, options=options).show()
+----+-------+
| id | name |
+----+-------+
| 1 | alice |
| 2 | N/A | <- expected NULL
| 3 | carol |
+----+-------+
Every entry point behaves the same way — the CsvReadOptions(null_regex=...)
constructor, with_null_regex(), read_csv(), register_csv(), and SQL:
ctx.sql("""
CREATE EXTERNAL TABLE t (id INT, name VARCHAR)
STORED AS CSV LOCATION 'probe.csv'
OPTIONS ('format.has_header' 'true', 'format.null_regex' '^(null|NULL|N/A)$')
""")
ctx.sql("select * from t").show() # same output, N/A not nulled
The SQL form does not reject the option, and giving an explicit schema does not
change anything, so this is not schema inference choosing Utf8.
When the matching value is in a numeric column, the read fails rather than
returning the wrong value:
p.write_text("id,value\n1,10\n2,N/A\n3,30\n")
ctx.sql("""CREATE EXTERNAL TABLE t2 (id INT, value BIGINT) STORED AS CSV
LOCATION 'probe.csv'
OPTIONS ('format.has_header' 'true', 'format.null_regex' '^(null|NULL|N/A)$')""")
ctx.sql("select * from t2").show()
DataFusion error: Arrow error: Parser error: Error while parsing value 'N/A' as
type 'Int64' for column 1 at line 2. Row data: '[2,N/A]'
This is the case that matters in practice: N/A, NULL and - placeholders
in otherwise numeric columns are the reason to reach for null_regex at all.
Expected behavior
A field matching null_regex is read as NULL, whatever the column's type.
Additional context
Why the existing coverage does not catch it: test_read_csv_with_options in
python/tests/test_context.py does set null_regex="[pP]+aris", but the only
Paris in its fixture is on the #Charlie;35;Paris line, which comment="#"
removes before the reader sees it. The None in that test's expected output
comes from truncated_rows=True on the Bob;25 row, not from null_regex. The
test pins that the option parses, which is what its comment says it is for — it
does not pin the behavior.
docs/source/user-guide/io/csv.md documents with_null_regex as "Treat these
as NULL", so the documented behavior and the actual behavior disagree.
Root cause (upstream). In datafusion-datasource-csv 55.0.0, the version
this repository pins, null_regex is applied during schema inference but never
to the reader that actually parses the rows:
src/file_format.rs:547sets it on thearrow::csv::reader::Formatused for
infer_schema.src/source.rs:187builds thecsv::ReaderBuilderthat reads the data, and
callswith_delimiter,with_header,with_quote,with_truncated_rows,
with_terminator,with_escapeandwith_comment— but not
with_null_regex.arrow_csv::reader::ReaderBuilder::with_null_regexexists
in arrow 59.2.0 and is simply never called.
CsvSource already holds the full CsvOptions, so self.options.null_regex is
in scope at that point — it looks like a few lines in builder(), mirroring the
escape and comment blocks just below it.
I could not find this reported in either this repository or apache/datafusion.
If you would rather track it upstream, I am happy to open it against
apache/datafusion and link it back here.
Found while working on #1728 / #1732, which is why the advanced-options example
there keeps with_null_regex set but puts its N/A in a string column, so the
example runs. Happy to take the fix if it turns out to be on this side.
- 主要語言
- Python
- 星號
- 605
- 分支
- 176
- 平均合併
- 1 天 23 小時
- 30 天內合併 PR
- 8
貢獻指南
這個儲存庫沒有索引到貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/datafusion-python 的其他 Issue
-
documentation
難度 2/5 1-3 小時 新手友好度 72/100
apache/datafusion-python#1726 ·
-
難度 2/5 半天 新手友好度 88/100
apache/datafusion-python#1691 ·
-
bug
難度 2/5 1-3 小時 新手友好度 78/100
apache/datafusion-python#1644 ·
-
enhancement
難度 5/5 一週以上 新手友好度 30/100
apache/datafusion-python#1737 ·
-
bug good first issue
難度 4/5 3-5 天 新手友好度 48/100
apache/datafusion-python#1728 ·
查看 apache/datafusion-python 的全部 Issue
相似的 Issue
-
documentation help wanted
難度 2/5 1-3 小時 新手友好度 90/100
-
難度 2/5 1-3 小時 新手友好度 90/100
simonw/sqlite-utils#872 ·
-
難度 2/5 1-3 小時 新手友好度 88/100
-
難度 2/5 1-3 小時 新手友好度 82/100
-
難度 2/5 1-3 小時 新手友好度 78/100