Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Add support for bucket expression to table scans

未关闭
#3,839 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
55/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
python
领域
databases

调研方向

从 issue 中描述的 table.scan 和 scan.to_arrow 入口点开始,然后跟踪用于 Arrow 分区裁剪的现有时间范围表达式处理逻辑。检查分区规范和行过滤器是如何应用的;当 bucket16 这样的 bucket 表达式能够在可能的情况下裁剪文件,并在无法裁剪时回退到行过滤时,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Feature Request / Improvement

For time partitioning, we can express time range expressions and they lead to partition pruning when planning an Arrow scan:

 scan = table.scan(
      row_filter=And(
          GreaterThanOrEqual("event_ts", start),
          LessThan("event_ts", end),
      )
  )

  arrow_table = scan.to_arrow()

It would be useful to be able to filter by other hidden partitioning transforms, such as buckets -- e.g. filtering on bucket[16](user_id) in {0, 1, 2, 3}, and getting partition pruning whenever possible based on the underlying table partitioning specs (falling back to filtering rows when files cannot be pruned).

主要语言
Python
星标
1.1k
派生
589
平均合并
2 天 2 小时
30 天内合并 PR
70

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/iceberg-python 的其他 Issue

查看 apache/iceberg-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。