Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

enhancement: stop admitted workloads from tolerating every node taint

Open
#599 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
go, kubernetes

Research direction

Start with pkg/worker/controllers/worker/node_flavor.go and the “eligibility by nodeLabels, not taints” section of docs/architecture/scheduling-chain.md. Compare the workload problem with the chart Deployment fix in issue #586, then verify the chosen behavior on the kind scenario described. Done requires an agreed design plus linked design, API, and documentation artifacts.

Written by the indexing model from the issue text.

Description

kind/enhancement

What would you like to be added:

Every ResourceFlavor the operator derives carries tolerations: [{operator: Exists}]
(pkg/worker/controllers/worker/node_flavor.go, the flavor spec; the rationale in the comment and in
docs/architecture/scheduling-chain.md, "eligibility by nodeLabels, not taints"). Kueue copies a
flavor's tolerations into every Pod it admits through that flavor, so each admitted workload tolerates
every node taint.

Replace the blanket toleration with one that still lets quota routing ignore the taints the operator
does not care about, but stops tolerating the taints that mean "do not run here now":

  • node.kubernetes.io/unschedulable (cordon, drain, autoscaler scale-down)
  • node.kubernetes.io/not-ready and node.kubernetes.io/unreachable with NoExecute, so the default
    300 s eviction applies again
  • the cluster-autoscaler and Karpenter disruption taints (ToBeDeletedByClusterAutoscaler,
    karpenter.sh/disrupted)

The shape of the replacement (an explicit allow-list of control-plane and accelerator taints, a
setting, or keeping Exists but excluding the lifecycle taints in the Pod webhook) is open.

Why is this needed:

  • Measured: on a local kind cluster whose control-plane node carries NoSchedule, a Workload admitted
    through a derived queue was placed by topology-aware scheduling on that control-plane node and
    reserved quota there. The blanket toleration makes TAS treat every taint as tolerated.
  • Read (not yet measured on a real accelerator node): because the Pod carries an empty-key Exists
    toleration, the DefaultTolerationSeconds admission plugin no longer adds the 300 s not-ready /
    unreachable NoExecute tolerations. A serving Pod on a node that goes unreachable is therefore
    never evicted and stays bound to the dead node until its Node object is deleted.
  • Inferred: a serving Pod can be placed on a node that is cordoned for maintenance or being drained,
    and a node being scaled in by an autoscaler keeps receiving new replicas.
  • The same failure shape was fixed for the chart's own Deployments (Kueue controller manager, NFD
    master, NFD gc) in #586; this is the workload side of it.

Where the need came from (tick one):

  • Blocked — something cannot be done today; name it, and who ran into it
  • Asked for — somebody outside this repository asked for it
  • Anticipated — nobody is blocked yet; this is expected to be needed

Found during the operator's own real-cluster and kind verification rounds, not reported by a user.
The people affected are cluster administrators who cordon, drain or autoscale accelerator nodes,
and anyone relying on Kubernetes' default eviction from failed nodes.

Completion requirements:

This enhancement requires the following artifacts:

  • Design doc
  • API change
  • Docs update

The artifacts should be linked in subsequent comments.

Dominant language
Go
Stars
4
Forks
6
Avg merge
1h 52m
Merged PRs (30d)
375

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from gpustack/gpustack-operator

All issues in gpustack/gpustack-operator

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.