ArrowScan to_table fails if the data is mixed between dict-encoded strings and plain strings.
維護者通常 1 天內回覆
@gyli 已經在處理了。
開始於 2026年10月8日。
評估
研究方向
從最小重現開始,檢查 ArrowScan.to_table 以及 issue 中所示的 to_record_batches 入口點。驗證如何合併混合了一般字串和字典編碼值的批次。完成標準是:重現過程在沒有 ArrowTypeError 的情況下完成,並回傳包含預期字串欄位的 Arrow 表格。
由索引模型根據 Issue 內容生成。
描述
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
We have recently updated our functions to call the pyiceberg table.append() function with dict encoded arrow tables. Now we have in our iceberg tables mixed data from before this change, (where our data still is stored as string) and after the change, where the data is stored as dict-encoded strings.
If we now call to_arrow() of a DataScan class, on this table we get this error:
pyarrow.lib.ArrowTypeError: Unable to merge: Field col has incompatible types: string vs dictionary<values=string, indices=int32, ordered=0>
Here is a minimal example that reproduces this error:
from pyiceberg.io.pyarrow import ArrowScan
from pyiceberg.table import ALWAYS_TRUE
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField
from pyiceberg.types import StringType
import pyarrow as pa
def create_scan_with_mixed_dict_encode_not_encode() -> ArrowScan:
schema = Schema(
NestedField(field_id=1, name="col", field_type=StringType(), required=False)
)
class FakeTableMetadata:
def schema(self) -> Schema:
return schema
scan = ArrowScan(table_metadata=FakeTableMetadata(),
io=object(),
projected_schema=schema,
row_filter=ALWAYS_TRUE)
def _batches_for_repro(self, _tasks):
str_values = pa.array(["a"], type=pa.string())
yield pa.record_batch([str_values], names=["col"])
yield pa.record_batch([str_values.dictionary_encode()], names=["col"])
ArrowScan.to_record_batches = _batches_for_repro
return scan
if __name__ == "__main__":
scan = create_scan_with_mixed_dict_encode_not_encode()
arrow_table = ArrowScan.to_table(scan, tasks=[])
I am happy to provide a bugfix PR, but I need a small guidance on the best approach.
One idea is to cast each batch in to_table to the arrow_schema. The more performant way is to check for each batch, if the schema is different. If they are different, then find the dict_encoded col and only cast that one to string.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- 主要語言
- Python
- 星號
- 1.2k
- 分支
- 618
- 平均合併
- 1 天 10 小時
- 30 天內合併 PR
- 71
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 沒有貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/iceberg-python 的其他 Issue
-
難度 2/5 1-3 小時 新手友好度 70/100
apache/iceberg-python#4093 ·
維護者通常 1 天內回覆
-
Replace `__slots__ = (field1,field2,...)` with `slots=True`可能已有人在做 @med9110 於 1 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 68/100
apache/iceberg-python#4086 · 3 則留言 ·
維護者通常 1 天內回覆
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's response可能已有人在做 @Soumo-git-hub 於 2 天前認領。 未關閉kind:bug
難度 2/5 1-3 小時 新手友好度 84/100
apache/iceberg-python#4073 · 1 則留言 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 70/100
apache/iceberg-python#4010 · 3 則留言 · 1 個 reaction ·
維護者通常 1 天內回覆
-
to_bytes silently rescales a Decimal with a negative scale可能已有人在做 @Rodrigo-Palma 於 22 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 78/100
apache/iceberg-python#3996 ·
維護者通常 1 天內回覆
查看 apache/iceberg-python 的全部 Issue
相似的 Issue
-
dependencies feature github_actions good first issue
難度 2/5 1-3 小時 新手友好度 62/100
wemake-services/wemake-django-template#3149 ·
維護者通常 1 天內回覆
-
upstream update
難度 2/5 1-3 小時 新手友好度 65/100
conan-io/conan-center-index#31142 ·
維護者通常 1 天內回覆
-
area:core bug
難度 2/5 1-3 小時 新手友好度 78/100
維護者通常 1 天內回覆
-
request-theme
難度 2/5 1 小時以內 新手友好度 70/100
LizardByte/ThemerrDB#8877 · 1 則留言 ·
維護者通常 1 天內回覆
-
area/install-update comp/gateway P0 sweeper:risk-compatibility type/bug
難度 2/5 1 小時以內 新手友好度 72/100
NousResearch/hermes-agent#135997 · 3 則留言 ·
維護者通常 1 天內回覆