Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Add a minimal micro-benchmark suite to validate filtering hot-path optimizations

未关闭
#690 3 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
35/100
Issue 类型
功能
描述清晰度
需要澄清
活跃度
冷清
技术栈
cpp

调研方向

首先检查现有的 CMake 和 GoogleTest 设置,然后跟踪 src/iceberg/**/ 下的过滤路径及其指标评估器。当项目具备一套已达成共识的最小基准测试套件、一个默认关闭的构建选项,以及一个隔离测量过滤步骤的基准测试时,这项工作就完成了。

由索引模型根据 Issue 内容生成。

描述

While reading the scan-planning filtering path, I found a small optimization in the metrics evaluators. Using it as a concrete example to raise a broader question about how to validate this kind of change.

Proposed change

The metrics evaluators run per data file. Each predicate currently calls expr->reference() repeatedly, and reference() returns a shared_ptr via shared_from_this() — an atomic refcount bump every time. The StrictMetricsEvaluator macro even discards a dynamic_cast result only to re-fetch the same reference:

- #define RETURN_IF_NOT_REFERENCE(expr)                                         \
-   if (auto ref = dynamic_cast<BoundReference*>(expr.get()); ref == nullptr) { \
-     return kRowsMightNotMatch;                                                \
-   }
+ #define BIND_REFERENCE_OR_RETURN(ref, expr)                            \
+   const auto* ref = dynamic_cast<const BoundReference*>((expr).get()); \
+   if (ref == nullptr) {                                                \
+     return kRowsMightNotMatch;                                         \
+   }

    Result<bool> IsNull(const std::shared_ptr<Bound>& expr) override {
-     RETURN_IF_NOT_REFERENCE(expr);
-     int32_t id = expr->reference()->field().field_id();
+     BIND_REFERENCE_OR_RETURN(ref, expr);
+     int32_t id = ref->field().field_id();
      ...

Reusing the cast result drops the repeated virtual reference() calls (and their atomic ops) across every predicate, with no behavior change.

Expected benefit

The win is on the CPU-bound filtering step, evaluated in isolation. Scan planning as a whole is IO-bound, so on an e2e scan this kind of change is almost certainly unmeasurable — which is exactly why it needs to be measured on the filtering step alone.

Which raises the question: do we need a benchmark suite?

This is exactly the kind of change that's hard to justify without one. The repo has no benchmark infrastructure today, only the gtest suite. A minimal benchmark on the filtering path would let us measure such changes on the CPU-bound step alone, rather than guessing or claiming a win against IO-dominated planning.

So before going further:

  1. How should we add it — Google Benchmark fetched the same way googletest already is, behind an off-by-default CMake option?
  2. Where should it live — a top-level benchmark/, or co-located under src/iceberg/**/?

I'm happy to put up a draft PR for a minimal suite + the filtering benchmark above once there's agreement on direction.

主要语言
C++
星标
223
派生
127
平均合并
1 天 13 小时
30 天内合并 PR
28

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/iceberg-cpp 的其他 Issue

查看 apache/iceberg-cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。