Optimize node collection indexes
还没有人认领这个 Issue。
评估
调研方向
先查看 issue 中的生产索引清单和大小查询,然后检查 kernelci-core/api/models.py 中的 Node.get_indexes()。在删除 id_1 之前验证 id 字段数量,删除或隐藏指定的索引,并在观察期结束后重新检查 $indexStats。剩余索引由代码管理且计划中的占用空间缩减已得到验证,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Background
The production node collection (35.6M docs) has 19 indexes totaling
~11.8 GB — 38% of all data on disk (storageSize 19 GB, indexSize
11.8 GB). Every node insert/update pays maintenance cost on all of them,
and they inflate physical backup size and restore time.
$indexStats from production (window: ~2 days since pod restart on
2026-07-06) shows roughly half of them unused.
Finding 1: no node indexes are managed by code
Node in kernelci-core does not implement get_indexes(), so
db.create_indexes() creates nothing for the node collection. All 18
non-_id indexes were created manually on the server with no record of
why. Whatever survives this cleanup should be added to
Node.get_indexes() (kernelci-core api/models.py) so the index set is
code-managed and reproducible.
Finding 2: redundant indexes — safe to drop now
data.kernel_revision.commit_1(12k ops) — fully covered by the
compound
data.kernel_revision.commit_1_data.kernel_revision.branch_1_data.kernel_revision.tree_1
(prefix rule). Drop the single, keep the compound.created_1_updated_-1(0 ops) —created_1is its prefix and is
in active use (likely the node purge job). Drop the compound, keep
created_1.id_1(0 ops) — the API mapsid↔_id; documents likely have
noidfield. Verify with
db.node.countDocuments({id: {$exists: true}})— if 0, drop.
Finding 3: zero-usage indexes — hide, observe, then drop
Zero ops in the sample window, but the window is too short to catch
weekly/monthly jobs. Hide them (planner stops using them, instantly
reversible with unhideIndex), observe 2–4 weeks, then drop what nobody
missed:
updated_1treeid_1(sparse)owner_1data.error_code_1result_1group_1data.kernel_revision.branch_1data.kernel_revision.tree_1
db.node.hideIndex("<name>")
Indexes confirmed in active use (keep)
_id_ (4.0M ops), parent_1 (2.3M), state_1 (712k), kind_1 (241k),
data.kernel_revision.commit_1_..._tree_1 compound,
processed_by_kcidb_bridge_1_updated_-1, name_1, created_1.
Expected impact
- Estimated 2–4 GB reduction in index footprint (more after observation
phase drops). - Reduced write amplification on every node insert/update.
- Smaller physical backups and faster restores.
Plan
- Capture per-index sizes:
db.node.aggregate([{$collStats: {storageStats: {}}}])→
indexSizes - Drop
data.kernel_revision.commit_1,created_1_updated_-1;
verify and dropid_1 - Hide the 8 zero-usage indexes, re-check
$indexStatsafter
2–4 weeks - Drop unused ones
- Add surviving indexes to
Node.get_indexes()in kernelci-core
Issue written with help of Codex AI assistant
- 主要语言
- Python
- 星标
- 10
- 派生
- 21
- 平均合并
- 22 分钟
- 30 天内合并 PR
- 1
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 没有贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
kernelci/kernelci-api 的其他 Issue
-
Components/Maestro/API/Local instance says GET /latest/ should return some JSON, but it doesn't可能重新可做 关联的 PR 已关闭且未合并。 未关闭documentation
难度 3/5 1-2 天 新手友好度 35/100
kernelci/kernelci-api#632 · 1 条评论 ·
-
难度 5/5 一周以上 新手友好度 20/100
kernelci/kernelci-api#629 ·
-
难度 4/5 3-5 天 新手友好度 35/100
kernelci/kernelci-api#608 ·
-
难度 3/5 1-2 天 新手友好度 35/100
kernelci/kernelci-api#597 · 2 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 48/100
kernelci/kernelci-api#582 ·
查看 kernelci/kernelci-api 的全部 Issue
相似的 Issue
-
changelog investigate
难度 2/5 1-3 小时 新手友好度 62/100
ramnes/notion-sdk-py#409 ·
-
good first issue help wanted
难度 2/5 1-3 小时 新手友好度 72/100
lindicaphxag-tech/kaggle#28 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
BSData/horus-heresy-3rd-edition#3211 ·
维护者通常 1 天内回复
-
bug needs-triage
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
-
Unreachable-proxy mount test depends on fixed port 9999可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭bug tests
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复