[Bug] Store ships rocksdb.total_memory_size of 32 GB, ignores the container memory limit, and Docker users cannot override it
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- docker, java, kubernetes
- Área
- backend, databases, infrastructure
Línea de trabajo
Start with application-pd.yml and AppConfig.java to trace how rocksdb.total_memory_size is selected, then read RaftRocksdbOptions.java and RocksDBOptions.java to understand the cache split. Inspect docker-entrypoint.sh and the deployment guide for configuration behavior and documented defaults. Done means the Store's RocksDB budget respects the process or container memory and Docker users can override it without breaking the documented configuration.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Bug Type (问题类型)
server status (启动/运行异常)
Before submit
- 我已经确认现有的 Issues 与 FAQ 中没有相同 / 重复问题 (I have confirmed and searched that there are no similar problems in the historical issue and documents)
Searched for total_memory_size, store OOM, SPRING_APPLICATION_JSON and application-pd.yml. #2847 contains total_memory_size: 64GB, but that is the reporter's own config pasted into a data-sync bug, not this problem.
Environment (环境信息)
- Code: apache/hugegraph master
176fb56dd(2026-10-02). All line references below are at that commit. - Backend: HStore, 3 Stores on k3s, cluster preset of the Helm chart in hugegraph/hugegraph#221 (chart
05f3d9e, PD/Store images of 2026-09-18), Store started with-Xmx1024m -XX:MaxDirectMemorySize=512m - Data Size: load of 5M vertices with a 1 KB text property each; the Stores died from vertex 2.36M on, holding about 1.0-1.3 GB of RocksDB data and 4.2 GB of raft log each
- Run and measurements by Sebastian Gruza, 2026-09-19: https://github.com/hugegraph/hugegraph/pull/221#issuecomment-5740338545
- All three Stores in that run log profiles
"pd", "default"andtotal_memory_size:32000000000at startup (store-1, lines 28 and 44-48). The comment's memory accounting assumes the key is unset and falls back to the heap size; these logs show it is set.
Expected & Actual behavior (期望与实际表现)
Expected: the Store's RocksDB budget fits inside the memory the process is given, or a Docker user can set it.
Actual: loading the data above OOM-killed all three Stores at a 4Gi container limit (Memory cgroup out of memory: Killed process (java) anon-rss:4154896kB), and they crash-looped afterwards. With the limit raised to 8Gi they stayed up, and store-1 measured 4.42 GiB anonymous RSS. The 8Gi limit only adds headroom; nothing bounds RocksDB below 32 GB.
Why nothing keeps RocksDB inside the limit, in code (the 4.42 GiB was not broken down per allocator, so the RocksDB share of it is not measured):
application-pd.yml:28-30setsrocksdb.total_memory_size: 32000000000. The shippedapplication.ymlincludes thepdprofile (application.yml:57-59). The copy inside the jar has the same value (hg-store-node/src/main/resources/application-pd.yml:37).- Because the key is set, the existing fallback in
AppConfig.java:104-107(useRuntime.maxMemory()when the key is missing or0) never runs. RaftRocksdbOptions.java:155-167splits the value into a write cache for theWriteBufferManagerand a block cache. With the defaultwrite_buffer_ratioof 0.66 (hg-store-rocksdb/.../RocksDBOptions.java:43-49) that is about 21.1 GB and 10.9 GB. Both are nativeLRUCaches, off-heap, so together they grow with data up to 32 GB and are not limited by the heap or the cgroup.RaftRocksdbOptions.java:69adds a separate 1 GBLRUCachefor the raft log storage on top.- The heap is limited separately.
hugegraph-store/Dockerfile:43sets-XX:MaxRAMPercentage=50, butstart-hugegraph-store.sh:145-152also passes-Xmxfromcalc_xmx(util.sh:179-197: 512 MB to 2 GB, half of free memory read from/proc/meminfo), and an explicit-Xmxtakes precedence. Neither value feeds intototal_memory_size. - Docker users cannot change it.
docker-entrypoint.sh:58-70buildsSPRING_APPLICATION_JSONfrom the fiveHG_STORE_*variables only and exports it, so aSPRING_APPLICATION_JSONpassed to the container is overwritten, and noHG_STORE_*variable maps to arocksdb.*key.
Possible workaround, untested: AppConfig.java:307-313 binds rocksdb as a map with @ConfigurationProperties(prefix = ""), so JAVA_OPTS="... -Drocksdb.total_memory_size=<bytes>" may reach the map key. Nobody has verified that it binds.
Fix direction (proposal):
- Derive the default from the memory the process actually has (cgroup limit, or the limit minus heap and direct memory) instead of a fixed 32 GB. Dropping the line from
application-pd.ymlalone would hand the budget to theRuntime.maxMemory()fallback, which sizes RocksDB equal to the heap; that is a smaller number but still a second heap-sized allocation on top of the first. - Add an
HG_STORE_ROCKSDB_TOTAL_MEMORY_SIZE(or similar) variable to the entrypoint's JSON, so container and Helm users can set it explicitly. hugegraph-store/docs/deployment-guide.md:348documents the 32 GB value and would need the same change.
Vertex/Edge example (问题点 / 边数据举例)
Not data specific. In the reported run the OOM came once each Store held about 1.0-1.3 GB of RocksDB data.
Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)
Not schema specific. The reported run used one vertex label with a 1 KB text property.
- Lenguaje dominante
- Java
- Estrellas
- 3.2k
- Forks
- 640
- Merge medio
- 2 d 9 h
- PR fusionados (30 d)
- 26
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/hugegraph
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
apache/hugegraph#3231 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[Bug] Prometheus metrics format bugQuizá libre de nuevo @cui2022 la tomó hace 59 días y no hay ningún pull request abierto. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 64/100
apache/hugegraph#3142 · 7 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
Los mantenedores suelen responder en 1 día
-
[Feature] Let the Server take the initial admin password without a properties-file round tripAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 45/100
Los mantenedores suelen responder en 1 día
-
[Bug] Basic auth decodes the credential as ASCII and splits on every colon: a non-ASCII password answers 401, a password with ':' answers 400Posiblemente ocupada @arshilkxwork la tomó hoy. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 35/100
apache/hugegraph#3284 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de apache/hugegraph
Issues similares
-
waiting-for-triage
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 84/100
ADORSYS-GIS/keycloak-oid4vp-plugin#221 ·
Los mantenedores suelen responder en 2 días
-
Upgrade to Spring Pulsar 2.0.8Abiertostatus: team-only type: dependency-upgrade
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
spring-projects/spring-boot#52099 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 67/100
tchiotludo/akhq#3307 · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
objectionary/jeo-maven-plugin#1885 ·
Los mantenedores suelen responder en 4 días