is_server_running marks a healthy job FAILED on libfabric `<warn>` lines (matches "unable to")
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 72/100
Research direction
Start in vec_inf/client/_utils.py at is_server_running and trace how stderr lines are classified before Application startup complete. Reproduce the libfabric case if possible, then add coverage for warning lines so a healthy server is not marked FAILED while preserving genuine fatal detection.
Written by the indexing model from the issue text.
Description
Version: vec-inf 0.9.0; the same logic is on main (vec_inf/client/_utils.py, is_server_running).
What happens: before Application startup complete. appears, any stderr line containing one of traceback, exception, fatal error, critical error, failed to, could not, unable to, error: returns (ModelStatus.FAILED, line). On an HPE Slingshot system (NCSA Delta) with libfabric logging enabled (FI_LOG_LEVEL=warn, common in Slingshot-enabled containers), the CXI provider writes warnings such as
libfabric:132:1787839803::cxi:ep_ctrl:cxip_ep_close():869<warn> gpua001.delta.ncsa.illinois.edu: Unable to free EP object -16 : Device or resource busy
to stderr during NCCL initialisation. The server continues and is ready about a minute later, but the status is already FAILED with that line as failed_reason, and a caller that acts on it (LLMHub's backend) cancels the healthy job. Observed on SLURM jobs 21500607 and 21500755 (vllm serve, tensor-parallel 4, single node); weights had loaded, KV cache was being allocated.
Suggestions (any one is enough):
- skip lines that carry an explicit warning marker (
<warn>,WARNING,warn:) before applying the fatal patterns; - or only report FAILED when the SLURM step has actually ended (or the ready signature is absent after a grace period), treating stderr matches as
pending_reasonuntil then; - or make the fatal patterns configurable in
environment.yaml.
Workaround in use: FI_LOG_PROV=none in the job environment, which silences the provider logs.
- Dominant language
- Python
- Stars
- 107
- Forks
- 14
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from VectorInstitute/vector-inference
-
New model request for Qwen3.8-27B on KillarneyMay be free again @mgsalem claimed this 47 days ago, and no pull request is open. Opennew model
VectorInstitute/vector-inference#303 · 1 assignee ·
-
New model request for Muse-Glimmer-30B on KillarneyMay be free again @mgsalem claimed this 47 days ago, and no pull request is open. Opennew model
VectorInstitute/vector-inference#302 · 1 assignee ·
-
Enable tool calling supportMay be free again @mgsalem claimed this 109 days ago, and no pull request is open. Openenhancement
VectorInstitute/vector-inference#266 · 1 comment · 1 assignee ·
-
new model
VectorInstitute/vector-inference#216 · 1 comment · 1 assignee ·
All issues in VectorInstitute/vector-inference
Similar issues
-
New InternshipOpennew_internship
Difficulty 1/5 Under an hour Newbie friendliness 70/100
-
[BUG] Reports tab: "Unban" button tooltip shows raw `{{ip}}` placeholder instead of the IP addressOpenbug javascript ui
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
bunkerity/bunkerweb#4001 · 1 comment ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
PedestrianDynamics/pyFDS-Evac#476 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
google/differential-privacy#516 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
adobe-fonts/source-serif#153 ·