kubernetes-sigs/cluster-api

Add maxRetry to RemediationStrategy in Machinedeployment

Aberta

#12.553 aberto em 30 de jul. de 2025

 (19 comentários) (2 reações) (0 responsável)Go (1.532 forks)auto 404
help wantedkind/featurepriority/backlogtriage/accepted

Métricas do repositório

Stars
 (4.267 estrelas)
Métricas de merge de PR
 (Métricas PR pendentes)

Description

What would you like to be added (User Story)?

Currently, Cluster API's KubeadmControlPlane (KCP) supports a remediationStrategy.maxretry field to control the number of times remediation is attempted before giving up. However, MachineDeployment lacks a similar capability.

Detailed Description

Proposal: Introduce a maxRetry field under MachineDeployment.spec.strategy.remediation that integrates with MachineHealthCheck. This field would define the maximum number of remediation attempts allowed for unhealthy machines associated with a MachineDeployment.

Use Case: In environments where aggressive or infinite remediation can lead to cascading failures or unnecessary resource churn, a retry limit helps provide guardrails for self-healing behaviour. Once maxRetry is reached, remediation would stop, and external signals (e.g., human intervention or alerting) can be relied upon.

Benefits:

  • Parity with KubeadmControlPlane's remediationStrategy.maxRetry.
  • Safer self-healing with better operational control.
  • Prevents runaway remediation loops in edge scenarios.
spec:
  strategy:
    remediation:
      maxRetry: <int>

Anything else you would like to add?

No response

Label(s) to be applied

/kind feature One or more /area label. See https://github.com/kubernetes-sigs/cluster-api/labels?q=area for the list of labels.

Guia do colaborador