[Bug]: Docker memory guard reads host RAM when no container limit is set — cgroup v2 "max" defeats get_container_memory_percent
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 72/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- docker, linux, python
- Lĩnh vực
- devops, infrastructure
Hướng nghiên cứu
Bắt đầu tại deploy/docker/utils.py quanh dòng 411 và kiểm tra cách get_container_memory_percent đọc memory.max cũng như xử lý lỗi phân tích cú pháp. Tái hiện bằng lệnh docker exec được cung cấp trong một container cgroup v2 không có giới hạn bộ nhớ. Được xem là hoàn tất khi trường hợp max không giới hạn sử dụng mẫu số tổng bộ nhớ host được tài liệu hóa mà không rơi vào fallback phần trăm host im lặng, đồng thời đường dẫn v1 hiện có vẫn nguyên vẹn.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.2 (observed on 0.9.0; the code path is identical on v0.9.2, main and develop)
Expected Behavior
get_container_memory_percent() reports the container's own usage against its own limit, as its docstring states ("cgroup v1/v2 aware"). With no limit set, it uses the host total as the denominator — the intent already expressed by the if limit > 1e18 branch.
Current Behavior
On cgroup v2 with no container memory limit, /sys/fs/cgroup/memory.max contains the literal string max. int("max") raises ValueError, the bare except swallows it, and the function returns psutil.virtual_memory().percent — the host's percentage.
deploy/docker/utils.py:411:
usage = int(usage_path.read_text())
limit = int(limit_path.read_text()) # ValueError on the string "max"
# Handle unlimited (v2: "max", v1: > 1e18)
if limit > 1e18: # only reachable on cgroup v1
import psutil
limit = psutil.virtual_memory().total
return (usage / limit) * 100
except:
import psutil
return psutil.virtual_memory().percent
The if limit > 1e18 branch is written for this case and its comment names the v2 "max" form, but on v2 the exception fires two lines earlier, so the branch never runs.
Effects:
memory_threshold_percentno longer guards the container. On a 16 GB host the 95% default resolves to ~15.2 GB used host-wide. In our case a worker grew to ~6 GB and the host-wide OOM killer fired at ~15.6 GB total RSS — the guard's remaining margin was ~400 MB.- The reading is coupled to unrelated containers. A memory-hungry neighbour can push it past the threshold, making crawl4ai refuse crawls while idle.
- The failure is silent. The bare
exceptlogs nothing, so a defeated guard looks like a working one.
Measured on a running container: 766 MiB in use (4.7% of the host) while the guard reported 50.3%.
Is this reproducible?
Yes
Inputs Causing the Bug
Docker server started with no memory limit — `docker run` without `-m`, or a
compose file without `deploy.resources.limits.memory` / `mem_limit` — on a
cgroup v2 host.
Steps to Reproduce
# 1. cgroup v2 reports no limit
docker exec crawl4ai cat /sys/fs/cgroup/memory.max
# max
# 2. what the guard sees
docker exec crawl4ai python3 -c \
"import sys; sys.path.insert(0,'/app'); from utils import get_container_memory_percent; print(get_container_memory_percent())"
# 50.3 <- host-wide usage
# 3. the container's real footprint at the same moment
docker stats --no-stream --format "{{.Name}} {{.MemUsage}}" crawl4ai
# 766MiB / 15.62GiB -> 4.7%
The shipped docker-compose.yml sets deploy.resources.limits.memory: 4G, so the documented compose path hides this. It surfaces with docker run and no -m, on Kubernetes pods without memory limits, and on PaaS UIs that set none by default.
Code snippets
Handle the v2 sentinel before int(), leaving the v1 path intact:
usage = int(usage_path.read_text())
raw_limit = limit_path.read_text().strip()
if raw_limit == "max": # cgroup v2: no limit
limit = psutil.virtual_memory().total
else:
limit = int(raw_limit)
if limit > 1e18: # cgroup v1: no limit
limit = psutil.virtual_memory().total
return (usage / limit) * 100
Narrowing except: to except (OSError, ValueError, ZeroDivisionError) with a log line would also make future parse failures visible rather than silently returning host figures — that silence is what made this hard to spot.
Happy to send a PR against develop if useful.
OS
Linux, cgroup v2 host (container: Debian bookworm)
Python version
3.12
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 58/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
Claiming namespace `apoint`Đang mởnamespace operations
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 82/100
EclipseFdn/open-vsx.org#13573 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
collective/icalendar#1854 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
rancher/rancher-ai-agent#412 ·
Maintainer thường phản hồi trong vòng 6 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
TUDelftGeodesy/DePSI#134 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
HenriquesLab/rxiv-maker#335 ·