kubernetes-sigs/cluster-api

Add maxRetry to RemediationStrategy in Machinedeployment

Ouverte

#12 553 ouverte le 30 juil. 2025

 (19 commentaires) (2 réactions) (0 personne assignée)Go (1 532 forks)auto 404
help wantedkind/featurepriority/backlogtriage/accepted

Métriques du dépôt

Stars
 (4 267 étoiles)
Métriques de merge PR
 (Métriques PR en attente)

Description

What would you like to be added (User Story)?

Currently, Cluster API's KubeadmControlPlane (KCP) supports a remediationStrategy.maxretry field to control the number of times remediation is attempted before giving up. However, MachineDeployment lacks a similar capability.

Detailed Description

Proposal: Introduce a maxRetry field under MachineDeployment.spec.strategy.remediation that integrates with MachineHealthCheck. This field would define the maximum number of remediation attempts allowed for unhealthy machines associated with a MachineDeployment.

Use Case: In environments where aggressive or infinite remediation can lead to cascading failures or unnecessary resource churn, a retry limit helps provide guardrails for self-healing behaviour. Once maxRetry is reached, remediation would stop, and external signals (e.g., human intervention or alerting) can be relied upon.

Benefits:

  • Parity with KubeadmControlPlane's remediationStrategy.maxRetry.
  • Safer self-healing with better operational control.
  • Prevents runaway remediation loops in edge scenarios.
spec:
  strategy:
    remediation:
      maxRetry: <int>

Anything else you would like to add?

No response

Label(s) to be applied

/kind feature One or more /area label. See https://github.com/kubernetes-sigs/cluster-api/labels?q=area for the list of labels.

Guide contributeur