enhancement: stop admitted workloads from tolerating every node taint
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- go, kubernetes
- Domain
- devops, infrastructure
Research direction
Start with pkg/worker/controllers/worker/node_flavor.go and the “eligibility by nodeLabels, not taints” section of docs/architecture/scheduling-chain.md. Compare the workload problem with the chart Deployment fix in issue #586, then verify the chosen behavior on the kind scenario described. Done requires an agreed design plus linked design, API, and documentation artifacts.
Written by the indexing model from the issue text.
Description
What would you like to be added:
Every ResourceFlavor the operator derives carries tolerations: [{operator: Exists}]
(pkg/worker/controllers/worker/node_flavor.go, the flavor spec; the rationale in the comment and in
docs/architecture/scheduling-chain.md, "eligibility by nodeLabels, not taints"). Kueue copies a
flavor's tolerations into every Pod it admits through that flavor, so each admitted workload tolerates
every node taint.
Replace the blanket toleration with one that still lets quota routing ignore the taints the operator
does not care about, but stops tolerating the taints that mean "do not run here now":
node.kubernetes.io/unschedulable(cordon, drain, autoscaler scale-down)node.kubernetes.io/not-readyandnode.kubernetes.io/unreachablewithNoExecute, so the default
300 s eviction applies again- the cluster-autoscaler and Karpenter disruption taints (
ToBeDeletedByClusterAutoscaler,
karpenter.sh/disrupted)
The shape of the replacement (an explicit allow-list of control-plane and accelerator taints, a
setting, or keeping Exists but excluding the lifecycle taints in the Pod webhook) is open.
Why is this needed:
- Measured: on a local kind cluster whose control-plane node carries
NoSchedule, a Workload admitted
through a derived queue was placed by topology-aware scheduling on that control-plane node and
reserved quota there. The blanket toleration makes TAS treat every taint as tolerated. - Read (not yet measured on a real accelerator node): because the Pod carries an empty-key
Exists
toleration, theDefaultTolerationSecondsadmission plugin no longer adds the 300 snot-ready/
unreachableNoExecutetolerations. A serving Pod on a node that goes unreachable is therefore
never evicted and stays bound to the dead node until itsNodeobject is deleted. - Inferred: a serving Pod can be placed on a node that is cordoned for maintenance or being drained,
and a node being scaled in by an autoscaler keeps receiving new replicas. - The same failure shape was fixed for the chart's own Deployments (Kueue controller manager, NFD
master, NFD gc) in #586; this is the workload side of it.
Where the need came from (tick one):
- Blocked — something cannot be done today; name it, and who ran into it
- Asked for — somebody outside this repository asked for it
- Anticipated — nobody is blocked yet; this is expected to be needed
Found during the operator's own real-cluster and kind verification rounds, not reported by a user.
The people affected are cluster administrators who cordon, drain or autoscale accelerator nodes,
and anyone relying on Kubernetes' default eviction from failed nodes.
Completion requirements:
This enhancement requires the following artifacts:
- Design doc
- API change
- Docs update
The artifacts should be linked in subsequent comments.
- Dominant language
- Go
- Stars
- 4
- Forks
- 6
- Avg merge
- 1h 52m
- Merged PRs (30d)
- 375
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from gpustack/gpustack-operator
-
enhancement: A recreated Devices ledger never restores the node's accelerator counting capacityOpenkind/enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
gpustack/gpustack-operator#146 ·
Maintainers usually reply within 1 day
-
area/worker kind/enhancement
Difficulty 5/5 Over a week Newbie friendliness 30/100
gpustack/gpustack-operator#693 ·
Maintainers usually reply within 1 day
-
todo
Difficulty 5/5 Over a week Newbie friendliness 35/100
gpustack/gpustack-operator#601 ·
Maintainers usually reply within 1 day
-
area/worker kind/bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
gpustack/gpustack-operator#547 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
gpustack/gpustack-operator#512 ·
Maintainers usually reply within 1 day
All issues in gpustack/gpustack-operator
Similar issues
-
cvss-severity:high devguard l3montree-cybersecurity/devguard/devguard pkg:golang/github.com/l3montree-dev/devguard risk:low state:open
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
l3montree-dev/devguard#3146 · 1 comment ·
Maintainers usually reply within 1 day
-
CLI
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
databricks/cli#6910 ·
Maintainers usually reply within 1 day
-
Table presenter appends a spurious ", ..." to the FIX column when all fix versions are already shownOpen
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
canonical/microceph#900 · 1 comment ·
Maintainers usually reply within 1 day
-
vulncheck or vulndb
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day