Add a minimal micro-benchmark suite to validate filtering hot-path optimizations
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 冷清
- 技术栈
- cpp
调研方向
首先检查现有的 CMake 和 GoogleTest 设置,然后跟踪 src/iceberg/**/ 下的过滤路径及其指标评估器。当项目具备一套已达成共识的最小基准测试套件、一个默认关闭的构建选项,以及一个隔离测量过滤步骤的基准测试时,这项工作就完成了。
由索引模型根据 Issue 内容生成。
描述
While reading the scan-planning filtering path, I found a small optimization in the metrics evaluators. Using it as a concrete example to raise a broader question about how to validate this kind of change.
Proposed change
The metrics evaluators run per data file. Each predicate currently calls expr->reference() repeatedly, and reference() returns a shared_ptr via shared_from_this() — an atomic refcount bump every time. The StrictMetricsEvaluator macro even discards a dynamic_cast result only to re-fetch the same reference:
- #define RETURN_IF_NOT_REFERENCE(expr) \
- if (auto ref = dynamic_cast<BoundReference*>(expr.get()); ref == nullptr) { \
- return kRowsMightNotMatch; \
- }
+ #define BIND_REFERENCE_OR_RETURN(ref, expr) \
+ const auto* ref = dynamic_cast<const BoundReference*>((expr).get()); \
+ if (ref == nullptr) { \
+ return kRowsMightNotMatch; \
+ }
Result<bool> IsNull(const std::shared_ptr<Bound>& expr) override {
- RETURN_IF_NOT_REFERENCE(expr);
- int32_t id = expr->reference()->field().field_id();
+ BIND_REFERENCE_OR_RETURN(ref, expr);
+ int32_t id = ref->field().field_id();
...
Reusing the cast result drops the repeated virtual reference() calls (and their atomic ops) across every predicate, with no behavior change.
Expected benefit
The win is on the CPU-bound filtering step, evaluated in isolation. Scan planning as a whole is IO-bound, so on an e2e scan this kind of change is almost certainly unmeasurable — which is exactly why it needs to be measured on the filtering step alone.
Which raises the question: do we need a benchmark suite?
This is exactly the kind of change that's hard to justify without one. The repo has no benchmark infrastructure today, only the gtest suite. A minimal benchmark on the filtering path would let us measure such changes on the CPU-bound step alone, rather than guessing or claiming a win against IO-dominated planning.
So before going further:
- How should we add it — Google Benchmark fetched the same way googletest already is, behind an off-by-default CMake option?
- Where should it live — a top-level
benchmark/, or co-located undersrc/iceberg/**/?
I'm happy to put up a draft PR for a minimal suite + the filtering benchmark above once there's agreement on direction.
- 主要语言
- C++
- 星标
- 223
- 派生
- 127
- 平均合并
- 1 天 13 小时
- 30 天内合并 PR
- 28
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/iceberg-cpp 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
apache/iceberg-cpp#973 ·
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 20/100
apache/iceberg-cpp#959 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 45/100
apache/iceberg-cpp#955 · 3 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 35/100
apache/iceberg-cpp#946 · 5 条评论 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 45/100
apache/iceberg-cpp#944 ·
维护者通常 1 天内回复
查看 apache/iceberg-cpp 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 92/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
cp-algorithms/cp-algorithms#1715 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
Icinga/icinga2#11058 · 1 条评论 ·
维护者通常 1 天内回复
-
status:needs-triage
难度 2/5 1-3 小时 新手友好度 88/100
PX4/PX4-Autopilot#28924 ·
维护者通常 1 天内回复
-
component: split-view platform: windows
难度 2/5 1-3 小时 新手友好度 74/100
zen-browser/desktop#15616 · 1 个 reaction ·
维护者通常 1 天内回复