Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

A file scan reads every file as UTF-8, whatever its Encoding field says

已关闭
#8,596 0 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 1 天内回复

@kz930 已经在做这个了。

开始于 2026年9月19日。

评估

这个 Issue 还没有评估数据。

描述

What happened?

File Scan offers an Encoding field. FileScanSourceOpDesc stores it in its own encoding property, but FileScanSourceOpExec decodes with fileEncoding, the property inherited from ScanSourceOpDesc. The class carries @JsonIgnoreProperties(Array("limit", "offset", "fileEncoding")), so fileEncoding never survives serialization into the executor and is always its default, UTF_8.

Choosing any other charset therefore changes nothing. A UTF-16 file comes back decoded as UTF-8 rather than as its text.

Expected: the executor decodes with the charset the Encoding field names.

How to reproduce?

Deserialize a File Scan descriptor carrying "encoding":"UTF_16", write it out the way getPhysicalOp does, and read it back the way FileScanSourceOpExec does. fileEncoding comes back UTF_8, and the encoding the user chose is the only place UTF-16 survives.

In the UI: upload a UTF-16 text file, drop a File Scan on it, set Encoding to UTF_16 and run. The rows hold the file's bytes read as UTF-8, not its lines.

Version/Branch

1.4.0-incubating-SNAPSHOT (main)

Commit Hash (Optional)

2ab8ee0f2

What browsers are you seeing the problem on?

No response

Relevant log output
JSON = {"operatorType":"FileScan","dummyPropertyList":[],"encoding":"UTF_16","extract":false,"outputFileName":false,"attributeType":"string","attributeName":"line",...}
exec reads desc.fileEncoding = UTF_8
主要语言
Scala
星标
316
派生
189
平均合并
4 天 21 小时
30 天内合并 PR
162

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/texera 的其他 Issue

查看 apache/texera 的全部 Issue

相似的 Issue

更多 Scala Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。