Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

ClickHouse: Cold runs preload data into memory before the timer

Offen
#941 1 Kommentar 5 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
45/100
Issue-Typ
Bug
Klarheit
Größtenteils klar
Aktivitätsstatus
Ruhig
Tech-Stack
shell, sql

Rechercherichtung

Lies zuerst die true-cold-Definition in README.md sowie die Startup-Kommentare zu ClickHouse/install. Verfolge, wie das Laden beim Startup und das Leeren des Caches im Benchmark-Ablauf dargestellt werden. Erledigt ist die Aufgabe, wenn die Richtlinie für Cold-Runs ausdrücklich festgelegt ist und Benchmark-Verhalten und Dokumentation übereinstimmen.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Commit 7c1f7a3291521738a7731af38c994debb1ea3b1e enables preloading of the primary key on ClickHouse startup.

Quoting clickhouse/install:

# Force synchronous startup loading so the cold timer doesn't catch
# work that should have been amortized into ./start.
#
# Two independent layers of laziness contribute to the cold-query floor:
#
# 1. async_load_databases (server-level): with the default 1, the server
#    binds its listen port and answers SELECT 1 before user databases
#    have finished loading. ./check passes, then the first query stalls
#    waiting for the part loader.
#
# 2. primary_key_lazy_load and columns_and_secondary_indices_sizes_lazy_calculation
#    (MergeTree-level): even after parts are loaded, the in-memory
#    primary key and the per-column .size streams are populated lazily
#    on first query. With ~25 parts × ~80 columns × multiple metadata
#    files per column that's >1.8k file opens on the first query path,
#    contributing several hundred ms even on local NVMe.
#
# Both are eager-load toggles, not caching shortcuts: the same I/O
# happens either way, just before query timing instead of during it.
# Together they brought Q40 cold from ~3 s to ~1.5 s on c6a.4xlarge.

README.md defines true cold runs differently:

2.a) True cold runs. Before each first run of each query, all [...] database caches (e.g. buffer pools) are cleared.

In the current configuration, ClickHouse pre-warms the primary key data at startup, which defeats the clearing of all database caches. The measured cold runs times aren't truly cold anymore. While size-calculations might be considered metadata, primary key data is definitively data that's also queried by the benchmark queries.

I'm not sure there's a good definition which data is fine to be cached for a true cold run and which not.
Two of my suggestions below:

  1. Include startup times in "true cold" runs. This effectively prevents cheating by preloading data at startup.
  2. Drop the cold run altogether. "Cold" is clearly different for different systems and does not compare performance apples-to-apples
Vorherrschende Sprache
Shell
Sterne
1.1k
Forks
321
Ø Merge
3 Std. 57 Min.
Gemergte PRs (30 T.)
607

Entwicklungsumgebung

  • Kein Dockerfile und keine Docker-Compose-Datei
  • Hat eine Pull-Request-Vorlage
  • Kein Beitragsleitfaden

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus ClickHouse/ClickBench

Alle Issues in ClickHouse/ClickBench

Ähnliche Issues

Weitere Issues zu Shell/Bash

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.