[Bug]: Docker memory guard reads host RAM when no container limit is set — cgroup v2 "max" defeats get_container_memory_percent
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 72/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 冷清
- 技术栈
- docker, linux, python
- 领域
- devops, infrastructure
调研方向
从 deploy/docker/utils.py 第 411 行附近开始,检查 get_container_memory_percent 如何读取 memory.max 以及如何处理解析错误。在没有内存限制的 cgroup v2 容器中,使用提供的 docker exec 命令重现该问题。完成的标准是:无上限的 max 情况使用文档规定的主机总内存作为分母,而不会落入静默的主机百分比回退,同时现有的 v1 路径保持不变。
由索引模型根据 Issue 内容生成。
描述
crawl4ai version
0.9.2 (observed on 0.9.0; the code path is identical on v0.9.2, main and develop)
Expected Behavior
get_container_memory_percent() reports the container's own usage against its own limit, as its docstring states ("cgroup v1/v2 aware"). With no limit set, it uses the host total as the denominator — the intent already expressed by the if limit > 1e18 branch.
Current Behavior
On cgroup v2 with no container memory limit, /sys/fs/cgroup/memory.max contains the literal string max. int("max") raises ValueError, the bare except swallows it, and the function returns psutil.virtual_memory().percent — the host's percentage.
deploy/docker/utils.py:411:
usage = int(usage_path.read_text())
limit = int(limit_path.read_text()) # ValueError on the string "max"
# Handle unlimited (v2: "max", v1: > 1e18)
if limit > 1e18: # only reachable on cgroup v1
import psutil
limit = psutil.virtual_memory().total
return (usage / limit) * 100
except:
import psutil
return psutil.virtual_memory().percent
The if limit > 1e18 branch is written for this case and its comment names the v2 "max" form, but on v2 the exception fires two lines earlier, so the branch never runs.
Effects:
memory_threshold_percentno longer guards the container. On a 16 GB host the 95% default resolves to ~15.2 GB used host-wide. In our case a worker grew to ~6 GB and the host-wide OOM killer fired at ~15.6 GB total RSS — the guard's remaining margin was ~400 MB.- The reading is coupled to unrelated containers. A memory-hungry neighbour can push it past the threshold, making crawl4ai refuse crawls while idle.
- The failure is silent. The bare
exceptlogs nothing, so a defeated guard looks like a working one.
Measured on a running container: 766 MiB in use (4.7% of the host) while the guard reported 50.3%.
Is this reproducible?
Yes
Inputs Causing the Bug
Docker server started with no memory limit — `docker run` without `-m`, or a
compose file without `deploy.resources.limits.memory` / `mem_limit` — on a
cgroup v2 host.
Steps to Reproduce
# 1. cgroup v2 reports no limit
docker exec crawl4ai cat /sys/fs/cgroup/memory.max
# max
# 2. what the guard sees
docker exec crawl4ai python3 -c \
"import sys; sys.path.insert(0,'/app'); from utils import get_container_memory_percent; print(get_container_memory_percent())"
# 50.3 <- host-wide usage
# 3. the container's real footprint at the same moment
docker stats --no-stream --format "{{.Name}} {{.MemUsage}}" crawl4ai
# 766MiB / 15.62GiB -> 4.7%
The shipped docker-compose.yml sets deploy.resources.limits.memory: 4G, so the documented compose path hides this. It surfaces with docker run and no -m, on Kubernetes pods without memory limits, and on PaaS UIs that set none by default.
Code snippets
Handle the v2 sentinel before int(), leaving the v1 path intact:
usage = int(usage_path.read_text())
raw_limit = limit_path.read_text().strip()
if raw_limit == "max": # cgroup v2: no limit
limit = psutil.virtual_memory().total
else:
limit = int(raw_limit)
if limit > 1e18: # cgroup v1: no limit
limit = psutil.virtual_memory().total
return (usage / limit) * 100
Narrowing except: to except (OSError, ValueError, ZeroDivisionError) with a log line would also make future parse failures visible rather than silently returning host figures — that silence is what made this hard to spot.
Happy to send a PR against develop if useful.
OS
Linux, cgroup v2 host (container: Debian bookworm)
Python version
3.12
- 主要语言
- Python
- 星标
- 84.5k
- 派生
- 8.7k
- 平均合并
- 3 天 9 小时
- 30 天内合并 PR
- 17
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
unclecode/crawl4ai 的其他 Issue
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run未关闭
难度 2/5 1-3 小时 新手友好度 78/100
unclecode/crawl4ai#2309 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 84/100
unclecode/crawl4ai#2147 · 3 条评论 ·
维护者通常 1 天内回复
-
🐞 Bug 🩺 Needs Triage
难度 4/5 3-5 天 新手友好度 55/100
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 35/100
维护者通常 1 天内回复
-
难度 3/5 1-2 天 新手友好度 58/100
维护者通常 1 天内回复
查看 unclecode/crawl4ai 的全部 Issue
相似的 Issue
-
namespace operations
难度 1/5 1 小时以内 新手友好度 82/100
EclipseFdn/open-vsx.org#13573 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
collective/icalendar#1854 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
rancher/rancher-ai-agent#412 ·
维护者通常 6 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
TUDelftGeodesy/DePSI#134 ·
-
难度 2/5 1-3 小时 新手友好度 88/100
HenriquesLab/rxiv-maker#335 ·