[Bug] HStore Server image (`hugegraph/server`) keeps crash diagnostics only inside the container, so a restart on Kubernetes loses them
I maintainer di solito rispondono entro 1 giorno
@byteayan ci sta già lavorando.
Dal 2/10/2026.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 52/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Ambito
- devops, infrastructure, observability
Direzione di ricerca
Start with Dockerfile-hstore and compare it with the standalone Dockerfile, then inspect log4j2.xml and hugegraph-server.sh at the referenced lines; docker-entrypoint.sh shows how JAVA_OPTS reaches the script. Check docker/README.md for the deployment documentation. Done means startup diagnostics reach Kubernetes stdout, WARN-level errors from the named loggers do too, JVM crash files use the logs directory, and that directory is documented for mounting.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Bug Type (问题类型)
server status (启动/运行异常)
Before submit
- I have confirmed and searched that there are no similar problems in the historical issue and documents
Searched open and closed issues and PRs for hs_err, ErrorFile, HeapDumpPath, additivity, STDOUT_MODE, Dockerfile-hstore, docker logs, kubectl logs and Starting HugeGraphServer failed. #2979 / #2980 wired the console appender and set STDOUT_MODE=true in hugegraph-server/Dockerfile (the standalone hugegraph/hugegraph image), but not in hugegraph-server/Dockerfile-hstore, which builds hugegraph/server (docker/README.md:240). #3203 hit the same wall during a cold-start failure but tracked the retry bug, not the logging.
Environment (环境信息)
- Server Version:
hugegraph/serverimages. 2026-09-10: publishedlatest(built 2026-09-08) and images built from master60c8803d(#3203). 2026-09-22: images built from master83ef9f3. All line numbers below are from master176fb56dd(2026-10-02). - Backend: hstore, 3 PD + 3 Store + 3 Server
- OS: Kubernetes (kind), HugeGraph Helm chart (#3218, earlier #3132)
- Data Size: empty (2026-09-10) or a few test writes (2026-09-22); the failure is at Server startup
Expected & Actual behavior (期望与实际表现)
Observed. In the 2026-09-10 run (#3203), kubectl logs --previous for a Server that exited during startup ended with these lines and held no stack trace:
Connecting to HugeGraphServer (http://0.0.0.0:8080/graphs)...........Starting HugeGraphServer failed
See /hugegraph-server/logs/hugegraph-server.log for HugeGraphServer log output.
(printed by util.sh:362 and start-hugegraph.sh:117). The second line is the non-STDOUT_MODE branch. The cause was in logs/hugegraph-server.log, which went away with the container; the trace in #3203 was only recovered by mounting a volume at /hugegraph-server/logs before the first boot. On 2026-09-22 one Server started during a helm upgrade exited once with Starting HugeGraphServer failed, and its cause was lost the same way.
Expected. After a Server container exits, its stdout (kubectl logs --previous) shows the error, and the JVM crash files sit in a documented directory an operator can mount.
Actual. Four things keep diagnostics inside the container filesystem.
-
The HStore image does not set
STDOUT_MODE.Dockerfile-hstoresets onlyJAVA_OPTSandHUGEGRAPH_HOME(:43-44), and the publishedhugegraph/server:latest(amd64 config, created 2026-10-01, read from the Docker Hub registry on 2026-10-02) has noSTDOUT_MODEin itsEnv. Without it,hugegraph-server.shsends the JVM's stdout and stderr,consoleappender included, tologs/hugegraph-server-stdout.log(:266-270), so nothing the JVM logs reacheskubectl logs. -
Even with
STDOUT_MODE=true, five loggers write only to the file.log4j2.xmlsetsLOG_PATHtologs(:21) and thefileappender writes${LOG_PATH}/hugegraph-server.log(:32). Since #2980 the root logger andorg.apache.hugegraphalso go toconsole(:106-109, :126-129).org.apache.hadoop,org.apache.zookeeper,com.alipay.sofa,io.nettyandorg.apache.commonsareadditivity="false"with onlyfile(:110-124), so their errors never reach stdout. Thefileappender also hasimmediateFlush="false"(:34), so a killed JVM can lose its last lines from the file as well. The audit loggers (:130-135) and the slow-query log (:136-138) are file-only too. They are separate streams and out of scope here. -
JVM crash files land inside the install directory.
hugegraph-server.shsetsLOGS="$TOP/logs"(:41), passes-XX:HeapDumpPath=${LOGS}(:124) and runscd "${TOP}"beforeexec java(:91). Nothing in the repository sets-XX:ErrorFile(git grep ErrorFileon176fb56ddreturns nothing), so HotSpot writeshs_err_pid<pid>.logto the working directory,/hugegraph-serverin the image. The image adds nothing here.JAVA_OPTS(Dockerfile-hstore:43) reaches the script through-j(docker-entrypoint.sh:243) and is appended on line 124. Line 124 sits inside theJAVA_OPTIONSdefault block (:118-129), so a user who setsJAVA_OPTIONSalso loses-XX:+HeapDumpOnOutOfMemoryError. -
The only persistence is an image
VOLUME.Dockerfile-hstorehasWORKDIR /hugegraph-server/(:46) andVOLUME /hugegraph-server(:78). Docker keeps that anonymous volume acrossdocker restart. Kubernetes does not turn an imageVOLUMEinto a pod volume. A restart starts a new container, andlogs/,hs_err_pid*.logand heap dumps go with the old one.
Fix direction.
- Set
STDOUT_MODE="true"inDockerfile-hstore, as #2980 did inDockerfile. - In
log4j2.xml, also send WARN and above from the five loggers on lines 110-124 toconsole(aconsoleref with a WARN threshold, or a second console appender). INFO from netty and hadoop stays in the file. Audit and slow-query stay file-only. - In
hugegraph-server.sh, add-XX:ErrorFile=${LOGS}/hs_err_pid%p.logand move-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=${LOGS}out of the default block, so both apply whateverJAVA_OPTIONSholds. - Document
/hugegraph-server/logsas the directory to mount on Kubernetes. If #3253 (for #3236) lands, these flags should follow itsLOGS_OVERRIDEinstead of$TOP/logs.
Related: #2979 / #2980, #3203, #3236 / #3253, #3211 (non-root image, which would chown the same log directory).
Vertex/Edge example (问题点 / 边数据举例)
Not applicable.
Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)
Not applicable.
- Lingua principale
- Java
- Stelle
- 3.2k
- Fork
- 641
- Merge medio
- 2g 9h
- PR unite (30g)
- 26
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/hugegraph
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
apache/hugegraph#3231 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug] Prometheus metrics format bugForse di nuovo libera @cui2022 l’ha presa 60 giorni fa e non c’è nessuna pull request aperta. Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 64/100
apache/hugegraph#3142 · 7 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
I maintainer di solito rispondono entro 1 giorno
-
[Feature] Let the Server take the initial admin password without a properties-file round tripAperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 45/100
I maintainer di solito rispondono entro 1 giorno
-
[Bug] Basic auth decodes the credential as ASCII and splits on every colon: a non-ASCII password answers 401, a password with ':' answers 400Forse già presa @arshilkxwork l’ha presa 1 giorno fa. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 35/100
apache/hugegraph#3284 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di apache/hugegraph
Issue simili
-
[Bug] AI unread message badge counts a batch of new bubbles as one messageForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
apache/rocketmq-dashboard#5784 ·
I maintainer di solito rispondono entro 3 giorni
-
[i18n] 安装实例完成后的成功提示未正确本地化Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
PCL-Community/PCL-CE#3658 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
apache/skywalking#14127 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
Team/Identity Server Core Type/Improvement U2
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
wso2/product-is#28553 ·
I maintainer di solito rispondono entro 1 giorno