enhancement: decide whether engine Pods get host-fabric access, and how it is asked for
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- docker, go, kubernetes
- Domain
- backend, cloud, infrastructure
Research direction
The issue is a design decision about whether engine Pods should inherit host-fabric access (like RDMA/EFA) from their members. Start by reading the linked issue #286 for context on measurements. Look at the ModelDeployment template and KVCacheBackend resource definitions in the codebase. Understand how the resolve webhook currently mutates resources. The decision involves security implications and API design, not a specific code change to implement.
Written by the indexing model from the issue text.
Description
The decision
On a host-fabric transport (RDMA or EFA), should the engine Pods that consume a pool be granted
the fabric access their members already have — and if so, how is it asked for?
This is the product question left over from #286, which is being closed because the measurement it
was waiting on has been taken. That measurement narrowed the question considerably but does not
answer it: what the access costs is now known; whether to grant it is not.
What is already settled, so it is not re-litigated here
The failure mode is a silent fallback, measured. Both engine-side render branches — the store
connector and the point-to-point connector — are told the fabric protocol, auto-discover the
topology of their own container, find zero HCAs, and install the TCP transport with no error, no
warning and no status condition. The deployment reaches Ready and serves. The object says EFA and
the bytes go over TCP. Verbatim readings are in #286.
The access needed for initialization is one extended resource request, not four things. A
five-cell ablation on hand-built Pods (a controlled experiment, not operator-rendered shapes) removed
one element of the member-side base at a time:
| removed | fabric initializes? |
|---|---|
| nothing (baseline) | yes |
hostNetwork |
yes |
| the device resource request | NO — loud, rc=-1 |
| the device-tree hostPath mount | yes (that device plugin injects the node itself) |
IPC_LOCK + SYS_RESOURCE |
yes |
Three boundaries travel with that table and must not be dropped when it is quoted:
- The mount's redundancy is a property of that vendor's device plugin, not of the protocol. The
cell that removed the request but kept the mount found the device node visible and
world-writable and the open still refused — the mount makes it seeable, the request makes it
openable. - Capabilities were removed at initialization only. The probe registers no buffers, and locked
memory limits bite at registration time. - The probe initializes a transport; it moves no bytes. So "hostNetwork is unnecessary" is
established for initialization, not for a data plane — whether fabric traffic egresses a pod
network namespace is untested.
What has to be decided
Whether. Granting this to engine Pods puts a privilege on tenant-adjacent workloads. The blast
radius is wider than the cluster-scoped member DaemonSet even when the grant is one extended
resource, because the workload that carries it is user-authored. The member-side rule — a privilege
is requested, never inferred — has no obvious engine-side spelling: there is no field on the
engine's owning object that asks for it.
If yes, how. Three shapes, ordered by API surface, carried over from #286 with their costs:
- A. A boolean on the ModelDeployment template, beside the existing
privileged. The user names
the privilege explicitly, which matches the member-side rule. Cost: it splits the fabric contract
in two — the backend saysEFA, the user must know to flip this on every engine consuming it, and
nothing refuses the mismatch. - B. Derive it from the referenced backend. The resolve webhook already fetches the
KVCacheBackendat mutate time, so a ModelDeployment binding a fabric backend could have the same
access rendered into its engine container. One declaration, both sides consistent. Cost: the
privilege becomes inferred from a binding rather than requested on the workload — the rule turns
into "binding a fabric backend is the request", which needs an admission-time guard to stay
honest. It also cannot stand alone for an engine consuming a remote store with no local binding. - C. A pool-level opt-in. The pool that admits fabric members marks itself fabric-capable and
engines admitted against it inherit the access. Cost: the placement contract (per-role
instanceType) becomes a security contract, which it was never meant to be.
A fourth answer is legitimate and is not the absence of a decision: members only, engines stay on
TCP. If that is the answer it has to be written as a rule, with the consequence stated — see the
sibling issue on pool-side reuse, which is what that rule costs.
Note that the field shape #286 assumed for option B no longer exists: the member's fabric device
resource is no longer a declared field but is derived from the transport protocol, so "read
deviceResourceName the same way the member rendering does" now means "derive it the same way".
What does NOT close this
- Rendering the access without deciding
whether. The ablation says what it costs, not that it
should be granted. - Documenting the gap. It is documented on both pages a reader would look for it. A deployment
that silently serves at TCP speed is not helped by a page the operator's author read. - A green run on a cluster where the engines do get the access. What has judgement power is what
happens to a deployment whose engines do not — today that is a healthy-looking deployment, which
is the thing to fix or to declare intended. - Choosing a shape without saying what refuses a mismatch. All three shapes admit a state where
the backend declares a fabric and an engine does not carry the access. Whichever is chosen has to
say whether that state is refused, reported, or allowed.
/kind enhancement
/area kv-cache
- Dominant language
- Go
- Stars
- 4
- Forks
- 6
- Avg merge
- 1h 52m
- Merged PRs (30d)
- 375
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from gpustack/gpustack-operator
-
enhancement: A recreated Devices ledger never restores the node's accelerator counting capacityOpenkind/enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
gpustack/gpustack-operator#146 ·
Maintainers usually reply within 1 day
-
area/worker kind/enhancement
Difficulty 5/5 Over a week Newbie friendliness 30/100
gpustack/gpustack-operator#693 ·
Maintainers usually reply within 1 day
-
todo
Difficulty 5/5 Over a week Newbie friendliness 35/100
gpustack/gpustack-operator#601 ·
Maintainers usually reply within 1 day
-
kind/enhancement
Difficulty 5/5 Over a week Newbie friendliness 35/100
gpustack/gpustack-operator#599 ·
Maintainers usually reply within 1 day
-
area/worker kind/bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
gpustack/gpustack-operator#547 ·
Maintainers usually reply within 1 day
All issues in gpustack/gpustack-operator
Similar issues
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
open-telemetry/opentelemetry-go-compile-instrumentation#1417 ·
Maintainers usually reply within 2 days
-
agent-research-finding agent-research-recommend chore ready-for-agent
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
jordansmall/spindrift#4068 · 1 comment ·
Maintainers usually reply within 1 day
-
Type/Task
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
OpenNSW/nsw-srilanka#537 ·
Maintainers usually reply within 1 day
-
security
Difficulty 2/5 1-2 days Newbie friendliness 62/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
Maintainers usually reply within 1 day