Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

Optimize node collection indexes

Offen
#706 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
45/100
Issue-Typ
Refactoring
Klarheit
Größtenteils klar
Aktivitätsstatus
Ruhig
Tech-Stack
mongodb, python
Bereich
databases

Rechercherichtung

Beginne mit dem Produktionsinventar der Indizes und der Größenabfrage im Issue und untersuche dann Node.get_indexes() in kernelci-core/api/models.py. Überprüfe die Anzahl der id-Felder, bevor du id_1 entfernst, entferne oder verberge die angegebenen Indizes und prüfe $indexStats nach dem Beobachtungszeitraum erneut. Als abgeschlossen gilt die Aufgabe, wenn die verbleibenden Indizes vom Code verwaltet werden und die geplante Reduzierung des Footprints verifiziert ist.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Background

The production node collection (35.6M docs) has 19 indexes totaling
~11.8 GB — 38% of all data on disk
(storageSize 19 GB, indexSize
11.8 GB). Every node insert/update pays maintenance cost on all of them,
and they inflate physical backup size and restore time.

$indexStats from production (window: ~2 days since pod restart on
2026-07-06) shows roughly half of them unused.

Finding 1: no node indexes are managed by code

Node in kernelci-core does not implement get_indexes(), so
db.create_indexes() creates nothing for the node collection. All 18
non-_id indexes were created manually on the server with no record of
why. Whatever survives this cleanup should be added to
Node.get_indexes() (kernelci-core api/models.py) so the index set is
code-managed and reproducible.

Finding 2: redundant indexes — safe to drop now

  • data.kernel_revision.commit_1 (12k ops) — fully covered by the
    compound
    data.kernel_revision.commit_1_data.kernel_revision.branch_1_data.kernel_revision.tree_1
    (prefix rule). Drop the single, keep the compound.
  • created_1_updated_-1 (0 ops) — created_1 is its prefix and is
    in active use (likely the node purge job). Drop the compound, keep
    created_1.
  • id_1 (0 ops) — the API maps id ↔ _id; documents likely have
    no id field. Verify with
    db.node.countDocuments({id: {$exists: true}}) — if 0, drop.

Finding 3: zero-usage indexes — hide, observe, then drop

Zero ops in the sample window, but the window is too short to catch
weekly/monthly jobs. Hide them (planner stops using them, instantly
reversible with unhideIndex), observe 2–4 weeks, then drop what nobody
missed:

  • updated_1
  • treeid_1 (sparse)
  • owner_1
  • data.error_code_1
  • result_1
  • group_1
  • data.kernel_revision.branch_1
  • data.kernel_revision.tree_1
db.node.hideIndex("<name>")

Indexes confirmed in active use (keep)

_id_ (4.0M ops), parent_1 (2.3M), state_1 (712k), kind_1 (241k),
data.kernel_revision.commit_1_..._tree_1 compound,
processed_by_kcidb_bridge_1_updated_-1, name_1, created_1.

Expected impact

  • Estimated 2–4 GB reduction in index footprint (more after observation
    phase drops).
  • Reduced write amplification on every node insert/update.
  • Smaller physical backups and faster restores.

Plan

  1. Capture per-index sizes:
    db.node.aggregate([{$collStats: {storageStats: {}}}]) →
    indexSizes
  2. Drop data.kernel_revision.commit_1, created_1_updated_-1;
    verify and drop id_1
  3. Hide the 8 zero-usage indexes, re-check $indexStats after
    2–4 weeks
  4. Drop unused ones
  5. Add surviving indexes to Node.get_indexes() in kernelci-core

Issue written with help of Codex AI assistant

Vorherrschende Sprache
Python
Sterne
10
Forks
21
Ø Merge
22 Min.
Gemergte PRs (30 T.)
1

Entwicklungsumgebung

  • Enthält ein Dockerfile oder eine Docker-Compose-Datei
  • Keine Pull-Request-Vorlage
  • Kein Beitragsleitfaden

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus kernelci/kernelci-api

Alle Issues in kernelci/kernelci-api

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.