Avro schema is re-converted to an Iceberg schema on every manifest read during scan planning

未关闭 适合新手
#3,662 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
78/100
Issue 类型
重构
描述清晰度
描述清楚
活跃度
冷清
技术栈
python
领域
performance

调研方向

从 pyiceberg/avro/file.py 中的 AvroFileHeader.get_schema() 开始,跟踪 avro_to_iceberg 转换。使用重复的 manifest 调用 scan().plan_files(),观察重复的工作,并确认转换结果会按 schema 字符串复用,同时随着 manifest 数量增加,规划性能得到提升。

由索引模型根据 Issue 内容生成。

描述

Feature Request / Improvement

AvroFileHeader.get_schema() (pyiceberg/avro/file.py) runs the full avro_to_iceberg conversion every time an
Avro file is opened.

The inefficiency

  • Every manifest under a spec embeds an identical Avro schema string.
  • So during scan planning, that same conversion is repeated once per manifest.
  • The cost grows with the manifest count, even though only a couple of distinct schemas are ever involved.

Proposed fix

  • The conversion depends only on the schema string, so the result can be cached (keyed on that string).

Measured impact

  • scan().plan_files() on an unpartitioned 150-manifest table: ~86 ms → ~48 ms (~1.8× faster).
  • The saving grows as the number of manifests increases.

I am willing to contribute for this improvement.

主要语言
Python
星标
1.1k
派生
589
平均合并
2 天 4 小时
30 天内合并 PR
72

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/iceberg-python 的其他 Issue

查看 apache/iceberg-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。