Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug] Store ships rocksdb.total_memory_size of 32 GB, ignores the container memory limit, and Docker users cannot override it

Aperta
#3,254 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
52/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
docker, java, kubernetes

Direzione di ricerca

Start with application-pd.yml and AppConfig.java to trace how rocksdb.total_memory_size is selected, then read RaftRocksdbOptions.java and RocksDBOptions.java to understand the cache split. Inspect docker-entrypoint.sh and the deployment guide for configuration behavior and documented defaults. Done means the Store's RocksDB budget respects the process or container memory and Docker users can override it without breaking the documented configuration.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Bug Type (问题类型)

server status (启动/运行异常)

Before submit
  • 我已经确认现有的 Issues 与 FAQ 中没有相同 / 重复问题 (I have confirmed and searched that there are no similar problems in the historical issue and documents)

Searched for total_memory_size, store OOM, SPRING_APPLICATION_JSON and application-pd.yml. #2847 contains total_memory_size: 64GB, but that is the reporter's own config pasted into a data-sync bug, not this problem.

Environment (环境信息)
  • Code: apache/hugegraph master 176fb56dd (2026-10-02). All line references below are at that commit.
  • Backend: HStore, 3 Stores on k3s, cluster preset of the Helm chart in hugegraph/hugegraph#221 (chart 05f3d9e, PD/Store images of 2026-09-18), Store started with -Xmx1024m -XX:MaxDirectMemorySize=512m
  • Data Size: load of 5M vertices with a 1 KB text property each; the Stores died from vertex 2.36M on, holding about 1.0-1.3 GB of RocksDB data and 4.2 GB of raft log each
  • Run and measurements by Sebastian Gruza, 2026-09-19: https://github.com/hugegraph/hugegraph/pull/221#issuecomment-5740338545
  • All three Stores in that run log profiles "pd", "default" and total_memory_size:32000000000 at startup (store-1, lines 28 and 44-48). The comment's memory accounting assumes the key is unset and falls back to the heap size; these logs show it is set.
Expected & Actual behavior (期望与实际表现)

Expected: the Store's RocksDB budget fits inside the memory the process is given, or a Docker user can set it.

Actual: loading the data above OOM-killed all three Stores at a 4Gi container limit (Memory cgroup out of memory: Killed process (java) anon-rss:4154896kB), and they crash-looped afterwards. With the limit raised to 8Gi they stayed up, and store-1 measured 4.42 GiB anonymous RSS. The 8Gi limit only adds headroom; nothing bounds RocksDB below 32 GB.

Why nothing keeps RocksDB inside the limit, in code (the 4.42 GiB was not broken down per allocator, so the RocksDB share of it is not measured):

  1. application-pd.yml:28-30 sets rocksdb.total_memory_size: 32000000000. The shipped application.yml includes the pd profile (application.yml:57-59). The copy inside the jar has the same value (hg-store-node/src/main/resources/application-pd.yml:37).
  2. Because the key is set, the existing fallback in AppConfig.java:104-107 (use Runtime.maxMemory() when the key is missing or 0) never runs.
  3. RaftRocksdbOptions.java:155-167 splits the value into a write cache for the WriteBufferManager and a block cache. With the default write_buffer_ratio of 0.66 (hg-store-rocksdb/.../RocksDBOptions.java:43-49) that is about 21.1 GB and 10.9 GB. Both are native LRUCaches, off-heap, so together they grow with data up to 32 GB and are not limited by the heap or the cgroup. RaftRocksdbOptions.java:69 adds a separate 1 GB LRUCache for the raft log storage on top.
  4. The heap is limited separately. hugegraph-store/Dockerfile:43 sets -XX:MaxRAMPercentage=50, but start-hugegraph-store.sh:145-152 also passes -Xmx from calc_xmx (util.sh:179-197: 512 MB to 2 GB, half of free memory read from /proc/meminfo), and an explicit -Xmx takes precedence. Neither value feeds into total_memory_size.
  5. Docker users cannot change it. docker-entrypoint.sh:58-70 builds SPRING_APPLICATION_JSON from the five HG_STORE_* variables only and exports it, so a SPRING_APPLICATION_JSON passed to the container is overwritten, and no HG_STORE_* variable maps to a rocksdb.* key.

Possible workaround, untested: AppConfig.java:307-313 binds rocksdb as a map with @ConfigurationProperties(prefix = ""), so JAVA_OPTS="... -Drocksdb.total_memory_size=<bytes>" may reach the map key. Nobody has verified that it binds.

Fix direction (proposal):

  • Derive the default from the memory the process actually has (cgroup limit, or the limit minus heap and direct memory) instead of a fixed 32 GB. Dropping the line from application-pd.yml alone would hand the budget to the Runtime.maxMemory() fallback, which sizes RocksDB equal to the heap; that is a smaller number but still a second heap-sized allocation on top of the first.
  • Add an HG_STORE_ROCKSDB_TOTAL_MEMORY_SIZE (or similar) variable to the entrypoint's JSON, so container and Helm users can set it explicitly.
  • hugegraph-store/docs/deployment-guide.md:348 documents the 32 GB value and would need the same change.
Vertex/Edge example (问题点 / 边数据举例)

Not data specific. In the reported run the OOM came once each Store held about 1.0-1.3 GB of RocksDB data.

Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)

Not schema specific. The reported run used one vertex label with a 1 KB text property.

Lingua principale
Java
Stelle
3.2k
Fork
640
Merge medio
2g 9h
PR unite (30g)
26

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/hugegraph

Tutte le issue di apache/hugegraph

Issue simili

Altre issue su Java

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.