llm-routing: runtime guard exits after start when the cgroup swap file is missing
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- helm, kubernetes, python
- Domain
- backend, infrastructure
Research direction
Start with charts/gguf-backend/files/runtime_guard.py (also referenced under deploy/helm/llm-routing/spark) and trace cgroup_memory() and the runtime startup order. Check how the guard handles missing cgroup files before Popen; done means it either refuses to start with a clear message or safely treats a missing swap file as zero, with the behavior verified for the reported case.
Written by the indexing model from the issue text.
Description
Describe the bug
charts/gguf-backend/files/runtime_guard.py (deploy/helm/llm-routing/spark; #2331 moves it under recipes/ without changing it) reads /sys/fs/cgroup/memory.swap.current in cgroup_memory(), but only inside the supervision loop, after it has started the runtime with Popen. On a host without cgroup v2 swap accounting, that file does not exist. FileNotFoundError then ends the guard, and its container, right after the RPC or model server starts.
Steps or code to reproduce bug
Run a runtime pod on a node where /sys/fs/cgroup/memory.swap.current is absent inside the container, for example with swap accounting disabled.
Expected behavior
Check that memory.current, memory.swap.current and memory.events exist before starting the runtime, and refuse to start with a clear message. Alternatively, treat a missing swap file as zero swap usage if that is safe.
Additional context
Introduced in #2322. Found during review of #2331.
- Dominant language
- Go
- Stars
- 229
- Forks
- 85
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 330
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/nvcf
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
good-first-issue
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/nvcf#1174 · 1 comment ·
Maintainers usually reply within 1 day
-
quic-go versions are skewed across the two ends of the same connectionMay be free again A pull request for this issue was closed without being merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Pointing two helm charts from nvcf-ncp-staging to the gated selfhosted-ga team - docs 0.6.1May be free again A pull request for this issue was closed without being merged. Openneeds-triage
Difficulty 1/5 Under an hour Newbie friendliness 78/100
Maintainers usually reply within 1 day
Similar issues
-
kind/bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 4 days
-
bug needs-acceptance
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
vllm-project/semantic-router#4744 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
jaegertracing/jaeger#9794 ·
Maintainers usually reply within 1 day
-
ScalingModifiers formula fails with "formula returned non-float result" when expression evaluates to an integerPossibly taken @Sarthak-Pandey claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
kedacore/keda#8270 · 1 comment ·
Maintainers usually reply within 1 day