Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

llm-routing: runtime guard exits after start when the cgroup swap file is missing

Open Beginner friendly
#2,334 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
68/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
helm, kubernetes, python

Research direction

Start with charts/gguf-backend/files/runtime_guard.py (also referenced under deploy/helm/llm-routing/spark) and trace cgroup_memory() and the runtime startup order. Check how the guard handles missing cgroup files before Popen; done means it either refuses to start with a clear message or safely treats a missing swap file as zero, with the behavior verified for the reported case.

Written by the indexing model from the issue text.

Description

bug llm-stack needs-triage

Describe the bug

charts/gguf-backend/files/runtime_guard.py (deploy/helm/llm-routing/spark; #2331 moves it under recipes/ without changing it) reads /sys/fs/cgroup/memory.swap.current in cgroup_memory(), but only inside the supervision loop, after it has started the runtime with Popen. On a host without cgroup v2 swap accounting, that file does not exist. FileNotFoundError then ends the guard, and its container, right after the RPC or model server starts.

Steps or code to reproduce bug

Run a runtime pod on a node where /sys/fs/cgroup/memory.swap.current is absent inside the container, for example with swap accounting disabled.

Expected behavior

Check that memory.current, memory.swap.current and memory.events exist before starting the runtime, and refuse to start with a clear message. Alternatively, treat a missing swap file as zero swap usage if that is safe.

Additional context

Introduced in #2322. Found during review of #2331.

Dominant language
Go
Stars
229
Forks
85
Avg merge
1d 6h
Merged PRs (30d)
330

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/nvcf

All issues in NVIDIA/nvcf

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.