Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: Docker memory guard reads host RAM when no container limit is set — cgroup v2 "max" defeats get_container_memory_percent

未关闭 适合新手
#2,123 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
72/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
冷清
技术栈
docker, linux, python

调研方向

从 deploy/docker/utils.py 第 411 行附近开始,检查 get_container_memory_percent 如何读取 memory.max 以及如何处理解析错误。在没有内存限制的 cgroup v2 容器中,使用提供的 docker exec 命令重现该问题。完成的标准是:无上限的 max 情况使用文档规定的主机总内存作为分母,而不会落入静默的主机百分比回退,同时现有的 v1 路径保持不变。

由索引模型根据 Issue 内容生成。

描述

crawl4ai version

0.9.2 (observed on 0.9.0; the code path is identical on v0.9.2, main and develop)

Expected Behavior

get_container_memory_percent() reports the container's own usage against its own limit, as its docstring states ("cgroup v1/v2 aware"). With no limit set, it uses the host total as the denominator — the intent already expressed by the if limit > 1e18 branch.

Current Behavior

On cgroup v2 with no container memory limit, /sys/fs/cgroup/memory.max contains the literal string max. int("max") raises ValueError, the bare except swallows it, and the function returns psutil.virtual_memory().percent — the host's percentage.

deploy/docker/utils.py:411:

usage = int(usage_path.read_text())
limit = int(limit_path.read_text())      # ValueError on the string "max"

# Handle unlimited (v2: "max", v1: > 1e18)
if limit > 1e18:                         # only reachable on cgroup v1
    import psutil
    limit = psutil.virtual_memory().total

return (usage / limit) * 100
except:
    import psutil
    return psutil.virtual_memory().percent

The if limit > 1e18 branch is written for this case and its comment names the v2 "max" form, but on v2 the exception fires two lines earlier, so the branch never runs.

Effects:

  1. memory_threshold_percent no longer guards the container. On a 16 GB host the 95% default resolves to ~15.2 GB used host-wide. In our case a worker grew to ~6 GB and the host-wide OOM killer fired at ~15.6 GB total RSS — the guard's remaining margin was ~400 MB.
  2. The reading is coupled to unrelated containers. A memory-hungry neighbour can push it past the threshold, making crawl4ai refuse crawls while idle.
  3. The failure is silent. The bare except logs nothing, so a defeated guard looks like a working one.

Measured on a running container: 766 MiB in use (4.7% of the host) while the guard reported 50.3%.

Is this reproducible?

Yes

Inputs Causing the Bug
Docker server started with no memory limit — `docker run` without `-m`, or a
compose file without `deploy.resources.limits.memory` / `mem_limit` — on a
cgroup v2 host.
Steps to Reproduce
# 1. cgroup v2 reports no limit
docker exec crawl4ai cat /sys/fs/cgroup/memory.max
# max

# 2. what the guard sees
docker exec crawl4ai python3 -c \
  "import sys; sys.path.insert(0,'/app'); from utils import get_container_memory_percent; print(get_container_memory_percent())"
# 50.3        <- host-wide usage

# 3. the container's real footprint at the same moment
docker stats --no-stream --format "{{.Name}} {{.MemUsage}}" crawl4ai
# 766MiB / 15.62GiB   -> 4.7%

The shipped docker-compose.yml sets deploy.resources.limits.memory: 4G, so the documented compose path hides this. It surfaces with docker run and no -m, on Kubernetes pods without memory limits, and on PaaS UIs that set none by default.

Code snippets

Handle the v2 sentinel before int(), leaving the v1 path intact:

usage = int(usage_path.read_text())
raw_limit = limit_path.read_text().strip()

if raw_limit == "max":                 # cgroup v2: no limit
    limit = psutil.virtual_memory().total
else:
    limit = int(raw_limit)
    if limit > 1e18:                   # cgroup v1: no limit
        limit = psutil.virtual_memory().total

return (usage / limit) * 100

Narrowing except: to except (OSError, ValueError, ZeroDivisionError) with a log line would also make future parse failures visible rather than silently returning host figures — that silence is what made this hard to spot.

Happy to send a PR against develop if useful.

OS

Linux, cgroup v2 host (container: Debian bookworm)

Python version

3.12

主要语言
Python
星标
84.5k
派生
8.7k
平均合并
3 天 9 小时
30 天内合并 PR
17

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

unclecode/crawl4ai 的其他 Issue

查看 unclecode/crawl4ai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。