doris-parquet-single fails before measurement on c6a.xlarge: FE cannot commit its 8 GiB initial heap
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Ambito
- databases, performance, testing-qa
Direzione di ricerca
Start with the doris-parquet-single entry and its benchmark.sh, then inspect the attached FE logs and result-validation artifacts to confirm the startup failure and missing measurements. Done means either recording the failed c6a.xlarge result in the intended format or agreeing on and documenting a clearly marked heap-override rerun.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Proposed title
doris-parquet-single fails before measurement on c6a.xlarge: FE cannot commit its 8 GiB initial heap
Proposed body
What happened
I reran doris-parquet-single from ClickBench commit 5a56398c975bfd9f328f544894bcb92533ed134c on c6a.xlarge to investigate the missing result reported in #1461. The benchmark exited with code 1 after 1,602 seconds and produced 0 of the expected 129 result rows.
The run did not reach data loading or query measurement. Doris BE started, but FE failed during JVM startup. ./check then timed out because nothing was listening on port 9030. This is the same machine type as the empty c6a.xlarge result in #1461.[1] ClickBench includes c6a.xlarge in its AMD matrix and describes it as the 8 GB "smallest cloud instance."[2][3]
The immediate failure is the fixed -Xms8192m: the JVM tries to commit 8,589,934,592 bytes on a host where Linux sees 8,107,167,744 bytes, before accounting for the OS, native JVM memory, or the co-located BE process.
Environment
- EC2:
c6a.xlarge(4 vCPU, 8 GiB according to AWS)[4] - Memory visible to Linux: 7,731 MiB (
8,107,167,744bytes) - OS: Ubuntu 24.04, Linux
7.0.0-1012-aws, x86_64 - ClickBench:
5a56398c975bfd9f328f544894bcb92533ed134c - Entry:
doris-parquet-single, unchanged from that ClickBench revision - Doris:
4.1.0-rc01 - Java: OpenJDK 17.0.20.1
- Swap after
doris-parquet-single/install: disabled
This was a custom Terraform-orchestrated rerun, not a ClickBench GitHub Actions run. It cloned the pinned ClickBench commit and invoked the entry's unmodified benchmark.sh; git status --porcelain was empty before execution. The outer runner created 16 GiB of swap and enabled earlyoom, matching ClickBench's cloud-init.sh.in.[8] The unmodified Doris install script then ran swapoff -a before startup. Neither earlyoom nor the kernel recorded a kill during this run.
Raw evidence
The release configuration starts FE with an 8 GiB initial and maximum heap:
JAVA_OPTS_FOR_JDK_17="... -Xmx8192m -Xms8192m ..."
This matches the Doris source at tag 4.1.0-rc01 (commit ca33ebcebae5ec9842ec6ede1ffc0c6f46023315).[5]
fe/log/fe.out:
StdoutLogger 2026-09-24 23:05:06,889 Using Java version 17
StdoutLogger 2026-09-24 23:05:06,944 ... -Xmx8192m -Xms8192m ...
OpenJDK 64-Bit Server VM warning: INFO: os::commit_memory(0x0000000600000000, 8589934592, 0) failed; error='Not enough space' (errno=12)
# There is insufficient memory for the Java Runtime Environment to continue.
# Native memory allocation (mmap) failed to map 8589934592 bytes. Error detail: committing reserved memory.
Final readiness failure:
bench: ./check did not succeed within 300s
bench: last ./check stderr was:
doris: SELECT 1 failed: ERROR 2003 (HY000): Can't connect to MySQL server on '127.0.0.1:9030' (111)
doris: BE process is running
Evidence files prepared for this issue:
doris-parquet-single.tar.gz: runner metadata, benchmark log, result validation, memory state, and a 23-file manifestdoris-fe-logs.tar.gz: the FE startup log containing the JVM allocation failurecustom-runner.txt: the outer-runner source used for this run
Doris documents 16 GB as the production FE minimum. Its deployment table separately lists 8 GB for FE and 16 GB for BE and recommends deploying them separately.[6] ClickBench runs both on one 8 GiB host.
The c6a.2xlarge artifact in #1461 completed all serial measurements, but reports concurrent_error_ratio: 0.997; it is not a healthy end-to-end result.[7] Because it also doubles the vCPU count, it only establishes that the larger configuration reached serial measurements. The memory threshold and the outcome under a smaller FE heap remain unmeasured.
Suggested next step
Do you want this recorded as a failed c6a.xlarge run, or should I add a clearly marked small-machine FE heap override and rerun it?
If the failure should be recorded, please point me to the intended result representation. If you prefer the heap override, I can provide that run separately from this untuned baseline.
Sources
[1] https://github.com/ClickHouse/ClickBench/pull/1461 — ClickBench PR #1461 c6a.xlarge result
[2] https://github.com/ClickHouse/ClickBench/blob/5a56398c975bfd9f328f544894bcb92533ed134c/.github/workflows/benchmark-pr.yml — ClickBench benchmark machine matrix
[3] https://github.com/ClickHouse/ClickBench/blob/5a56398c975bfd9f328f544894bcb92533ed134c/CHANGELOG.md — ClickBench small machine rationale
[4] https://aws.amazon.com/ec2/instance-types/c6a — AWS C6a instance specifications
[5] https://github.com/apache/doris/blob/ca33ebcebae5ec9842ec6ede1ffc0c6f46023315/conf/fe.conf — Apache Doris 4.1.0-rc01 FE configuration
[6] https://doris.apache.org/docs/4.x/install/preparation/env-checking — Apache Doris hardware and software environment check
[7] https://github.com/ClickHouse/ClickBench/blob/5a56398c975bfd9f328f544894bcb92533ed134c/doris-parquet-single/results/20260819/c6a.2xlarge.json — Doris c6a.2xlarge ClickBench result
[8] https://github.com/ClickHouse/ClickBench/blob/5a56398c975bfd9f328f544894bcb92533ed134c/cloud-init.sh.in — ClickBench cloud-init swap and earlyoom setup
Files
custom-runner.txt
doris-fe-logs.tar.gz
doris-parquet-single.tar.gz
evidence-excerpts.txt
- Lingua principale
- Shell
- Stelle
- 1.1k
- Fork
- 315
- Merge medio
- 3h 43m
- PR unite (30g)
- 602
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di ClickHouse/ClickBench
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 55/100
ClickHouse/ClickBench#2146 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
ClickBench+Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
ClickHouse/ClickBench#1488 ·
I maintainer di solito rispondono entro 1 giorno
-
close in a month if not active help wanted
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
ClickHouse/ClickBench#1459 ·
I maintainer di solito rispondono entro 1 giorno
-
close in a month if not active
Difficoltà 3/5 1-2 giorni Idoneità per principianti 58/100
ClickHouse/ClickBench#1322 ·
I maintainer di solito rispondono entro 1 giorno
-
Add Apache KuduAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 25/100
ClickHouse/ClickBench#1090 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di ClickHouse/ClickBench
Issue simili
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
alunduil/alunduil-infrastructure#629 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
duckdb/duckdb-skills#19 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
YosysHQ/oss-cad-suite-build#216 ·