Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

is_server_running marks a healthy job FAILED on libfabric `<warn>` lines (matches "unable to")

Open
#308 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
72/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Start in vec_inf/client/_utils.py at is_server_running and trace how stderr lines are classified before Application startup complete. Reproduce the libfabric case if possible, then add coverage for warning lines so a healthy server is not marked FAILED while preserving genuine fatal detection.

Written by the indexing model from the issue text.

Description

Version: vec-inf 0.9.0; the same logic is on main (vec_inf/client/_utils.py, is_server_running).

What happens: before Application startup complete. appears, any stderr line containing one of traceback, exception, fatal error, critical error, failed to, could not, unable to, error: returns (ModelStatus.FAILED, line). On an HPE Slingshot system (NCSA Delta) with libfabric logging enabled (FI_LOG_LEVEL=warn, common in Slingshot-enabled containers), the CXI provider writes warnings such as

libfabric:132:1787839803::cxi:ep_ctrl:cxip_ep_close():869<warn> gpua001.delta.ncsa.illinois.edu: Unable to free EP object -16 : Device or resource busy

to stderr during NCCL initialisation. The server continues and is ready about a minute later, but the status is already FAILED with that line as failed_reason, and a caller that acts on it (LLMHub's backend) cancels the healthy job. Observed on SLURM jobs 21500607 and 21500755 (vllm serve, tensor-parallel 4, single node); weights had loaded, KV cache was being allocated.

Suggestions (any one is enough):

  • skip lines that carry an explicit warning marker (<warn>, WARNING, warn:) before applying the fatal patterns;
  • or only report FAILED when the SLURM step has actually ended (or the ready signature is absent after a grace period), treating stderr matches as pending_reason until then;
  • or make the fatal patterns configurable in environment.yaml.

Workaround in use: FI_LOG_PROV=none in the job environment, which silences the provider logs.

Dominant language
Python
Stars
107
Forks
14
PR merge metrics
No merged PRs in 30d

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from VectorInstitute/vector-inference

All issues in VectorInstitute/vector-inference

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.