[Bug] Store ships rocksdb.total_memory_size of 32 GB, ignores the container memory limit, and Docker users cannot override it
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 52/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- docker, java, kubernetes
- Lĩnh vực
- backend, databases, infrastructure
Hướng nghiên cứu
Start with application-pd.yml and AppConfig.java to trace how rocksdb.total_memory_size is selected, then read RaftRocksdbOptions.java and RocksDBOptions.java to understand the cache split. Inspect docker-entrypoint.sh and the deployment guide for configuration behavior and documented defaults. Done means the Store's RocksDB budget respects the process or container memory and Docker users can override it without breaking the documented configuration.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Bug Type (问题类型)
server status (启动/运行异常)
Before submit
- 我已经确认现有的 Issues 与 FAQ 中没有相同 / 重复问题 (I have confirmed and searched that there are no similar problems in the historical issue and documents)
Searched for total_memory_size, store OOM, SPRING_APPLICATION_JSON and application-pd.yml. #2847 contains total_memory_size: 64GB, but that is the reporter's own config pasted into a data-sync bug, not this problem.
Environment (环境信息)
- Code: apache/hugegraph master
176fb56dd(2026-10-02). All line references below are at that commit. - Backend: HStore, 3 Stores on k3s, cluster preset of the Helm chart in hugegraph/hugegraph#221 (chart
05f3d9e, PD/Store images of 2026-09-18), Store started with-Xmx1024m -XX:MaxDirectMemorySize=512m - Data Size: load of 5M vertices with a 1 KB text property each; the Stores died from vertex 2.36M on, holding about 1.0-1.3 GB of RocksDB data and 4.2 GB of raft log each
- Run and measurements by Sebastian Gruza, 2026-09-19: https://github.com/hugegraph/hugegraph/pull/221#issuecomment-5740338545
- All three Stores in that run log profiles
"pd", "default"andtotal_memory_size:32000000000at startup (store-1, lines 28 and 44-48). The comment's memory accounting assumes the key is unset and falls back to the heap size; these logs show it is set.
Expected & Actual behavior (期望与实际表现)
Expected: the Store's RocksDB budget fits inside the memory the process is given, or a Docker user can set it.
Actual: loading the data above OOM-killed all three Stores at a 4Gi container limit (Memory cgroup out of memory: Killed process (java) anon-rss:4154896kB), and they crash-looped afterwards. With the limit raised to 8Gi they stayed up, and store-1 measured 4.42 GiB anonymous RSS. The 8Gi limit only adds headroom; nothing bounds RocksDB below 32 GB.
Why nothing keeps RocksDB inside the limit, in code (the 4.42 GiB was not broken down per allocator, so the RocksDB share of it is not measured):
application-pd.yml:28-30setsrocksdb.total_memory_size: 32000000000. The shippedapplication.ymlincludes thepdprofile (application.yml:57-59). The copy inside the jar has the same value (hg-store-node/src/main/resources/application-pd.yml:37).- Because the key is set, the existing fallback in
AppConfig.java:104-107(useRuntime.maxMemory()when the key is missing or0) never runs. RaftRocksdbOptions.java:155-167splits the value into a write cache for theWriteBufferManagerand a block cache. With the defaultwrite_buffer_ratioof 0.66 (hg-store-rocksdb/.../RocksDBOptions.java:43-49) that is about 21.1 GB and 10.9 GB. Both are nativeLRUCaches, off-heap, so together they grow with data up to 32 GB and are not limited by the heap or the cgroup.RaftRocksdbOptions.java:69adds a separate 1 GBLRUCachefor the raft log storage on top.- The heap is limited separately.
hugegraph-store/Dockerfile:43sets-XX:MaxRAMPercentage=50, butstart-hugegraph-store.sh:145-152also passes-Xmxfromcalc_xmx(util.sh:179-197: 512 MB to 2 GB, half of free memory read from/proc/meminfo), and an explicit-Xmxtakes precedence. Neither value feeds intototal_memory_size. - Docker users cannot change it.
docker-entrypoint.sh:58-70buildsSPRING_APPLICATION_JSONfrom the fiveHG_STORE_*variables only and exports it, so aSPRING_APPLICATION_JSONpassed to the container is overwritten, and noHG_STORE_*variable maps to arocksdb.*key.
Possible workaround, untested: AppConfig.java:307-313 binds rocksdb as a map with @ConfigurationProperties(prefix = ""), so JAVA_OPTS="... -Drocksdb.total_memory_size=<bytes>" may reach the map key. Nobody has verified that it binds.
Fix direction (proposal):
- Derive the default from the memory the process actually has (cgroup limit, or the limit minus heap and direct memory) instead of a fixed 32 GB. Dropping the line from
application-pd.ymlalone would hand the budget to theRuntime.maxMemory()fallback, which sizes RocksDB equal to the heap; that is a smaller number but still a second heap-sized allocation on top of the first. - Add an
HG_STORE_ROCKSDB_TOTAL_MEMORY_SIZE(or similar) variable to the entrypoint's JSON, so container and Helm users can set it explicitly. hugegraph-store/docs/deployment-guide.md:348documents the 32 GB value and would need the same change.
Vertex/Edge example (问题点 / 边数据举例)
Not data specific. In the reported run the OOM came once each Store held about 1.0-1.3 GB of RocksDB data.
Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)
Not schema specific. The reported run used one vertex label with a 1 KB text property.
- Ngôn ngữ chính
- Java
- Star
- 3.2k
- Fork
- 640
- Merge trung bình
- 2 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 26
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/hugegraph
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
apache/hugegraph#3231 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug] Prometheus metrics format bugCó thể làm lại được @cui2022 đã nhận 58 ngày trước và không có pull request nào đang mở. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 64/100
apache/hugegraph#3142 · 7 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
Maintainer thường phản hồi trong vòng 1 ngày
-
[Feature] Let the Server take the initial admin password without a properties-file round tripĐang mở
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 45/100
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug] Basic auth decodes the credential as ASCII and splits on every colon: a non-ASCII password answers 401, a password with ':' answers 400Có thể đã có người làm @arshilkxwork đã nhận hôm nay. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 35/100
apache/hugegraph#3284 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của apache/hugegraph
Issue tương tự
-
Mend: dependency security vulnerability
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
opfab/operatorfabric-core#10653 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filtersCó thể đã có người làm @innoprej đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Broken links in the docsĐang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 78/100
salesforce/multicloudj#667 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 2 ngày