kubernetes-sigs/cluster-api

Stale `HealthCheckSucceeded=False` condition not cleared after node recovery when using external remediation

開放

#13,822 建立於 2026年6月18日

 (13 則留言) (0 個反應) (1 位負責人)Go (1,532 個分叉)auto 404
area/machinehealthcheckhelp wantedkind/bugpriority/important-soontriage/accepted

倉庫指標

星標
 (4,267 顆星)
PR 合併指標
 (PR 指標待抓取)

描述

We have encountered an issue where machines end up with a stale HealthCheckSucceeded=False condition that never gets cleared, even though the nodes have fully recovered and are running fine.

This has been observed in two scenarios, both using a MachineHealthCheck with a remediationTemplate:

  1. "No heartbeat for 5 minutes": Node briefly reported Ready=Unknown, the health check fired, the node recovered on its own, but HealthCheckSucceeded=False persists.
  2. "Exceeded start-up time of 30m": Node eventually came up healthy after exceeding the configured startup timeout, but the failed condition was never removed.

After a node recovers from such a transient health check failure:

  • NodeHealthy=True
  • NodeReady=True
  • No active remediation objects exist
  • HealthCheckSucceeded=False: this was not cleared
  • Available=False: machine appears unavailable when it is not

CAPI version used: v1.13.1

Label(s) to be applied

/kind bug /area machinehealthcheck

貢獻者指南