feat: Add liveness probes to detect stuck operators and restart instead of hanging silently
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 48/100
- Tipo de issue
- Funcionalidade
- Clareza
- Razoavelmente clara
- Status de atividade
- Pouca atividade
- Stack de tecnologia
- kubernetes, rust
- Domínio
- devops, observability
Direção de pesquisa
Comece localizando o código relacionado à saúde de operator-rs e os charts do Kubernetes para os deployments do operador; a issue diz que o código de probe existente só constrói probes para product-pod. Defina o comportamento de /livez em torno da alcançabilidade do API-server e de sua métrica, e então conecte probes de liveness menos rigorosas aos charts. Está concluído quando operadores travados puderem ser observados por meio de falhas nas probes, contagens de reinicialização e da métrica de health, sem reagir a quedas normais de watch-stream.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Observation
While testing demos on a managed Kubernetes cluster I hit a window where connections from inside the cluster to the API server kept hanging. Access from outside (kubectl on my laptop) was fine the whole time, and the kubelets were also fine - it was just the pod network path that was affected.
We can't change the infrastructure, but the way our operators behaved during and after that window is our problem, and we can fix that:
- The operators just logged
failed to start watching object: ... client error (Connect) ... deadline has elapsedrepeatedly. While that's going on the operator is blind. I created an AirflowCluster during the window and it got nothing - no status, no events, no resources - for half an hour. - Any pod being started that needs a secret-operator or listener-operator volume will fail to mount it while this is happening - existing pods keep running.
- The pods sit there Running and Ready, but there are no events or restarts. You only find out by reading the operator logs.
- They don't recover on their own: after the network was demonstrably healthy again (a fresh test pod connected 10 out of 10 times), the stuck operators kept failing for minutes, and only came back after a
rollout restart. So a short network blip can leave operators dead until someone notices.
N.B.: watch-stream drops (peer closed connection without sending TLS close_notify) occur continually (also while this was happening), but these are fine as clients can reconnect. Whatever health check we build must trigger on "can't open connections at all", not on a quiet watch stream, otherwise it'll restart operators on every idle cluster.
Where we are today
- None of our operator deployments have liveness or readiness probes.
- operator-rs has no health endpoint either - the only probe code in there is for building probes on the product pods.
So a stuck operator is invisible and stays stuck until manually restarted.
Proposal
- Provide a
/livezendpoint in operator-rs (so that all operators get it) that actually checks the API server is reachable and expose it as a metric - Add liveness probes to the charts (relaxed enough that a brief blip doesn't cause restart churn)
N.B.: a liveness restart might not even cure the stuck state (in our incident only replacing the pods provably did - the pod keeps its IP and network namespace across a container restart). The main win is that a stuck operator becomes visible - restart counts, probe failures, an alertable metric - instead of silently ignoring the cluster for half an hour.
To cover the case where a container restart isn't enough, we could add a small watchdog as a second step: something running on the host network (which stayed healthy throughout the incident) that checks the operators' /livez endpoints and deletes stuck pods so they get properly replaced. That's the guaranteed fix - pod replacement is what actually worked - but it only makes sense once the /livez endpoint from this issue exists, so it could be a follow-up issue or a second acceptance criterion here, whichever fits better.
See related issue: https://github.com/stackabletech/issues/issues/746
- Linguagem predominante
- Sem dados de linguagem
- Estrelas
- 2
- Forks
- 0
- Métricas de merge de PRs
- Nenhum PR com merge em 30d
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de stackabletech/issues
-
Metadata store: MVPAberta
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 25/100
stackabletech/issues#892 ·
-
Metadata StoreAberta
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 25/100
stackabletech/issues#891 · 1 comentário · 1 reação ·
-
Release Retro 26.11.0Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 50/100
stackabletech/issues#890 ·
-
tracking: SDP Release 26.11.0Talvez já em andamento @siegfriedweber assumiu há 13 dias. Abertaepic
stackabletech/issues#889 · 2 responsáveis ·
-
Kafka: Add support for managing Kafka topics in a declarative wayTalvez já em andamento @labrenbe assumiu há 19 dias. Aberta
stackabletech/issues#888 · 1 comentário · 1 responsável ·
Todas as issues de stackabletech/issues
Issues semelhantes
-
bot-found bug priority: P3
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
madenvel/KalinkaPlayer#179 ·
-
Dificuldade 1/5 1-3 horas Facilidade para iniciantes 94/100
makeplane/helm-charts#332 ·
Mantenedores costumam responder em até 1 dia
-
bug needs triage
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
microsoft/fabric-cicd#1141 · 1 comentário ·
-
[aw] Upgrade availableAbertaagentic-workflows
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
githubnext/gh-aw-workshop#3933 ·
Mantenedores costumam responder em até 2 dias
-
Remove CAAPFAbertakind/chore kind/cleanup needs-area
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 86/100
rancher/turtles#2848 · 3 comentários ·
Mantenedores costumam responder em até 1 dia