Make probe paths configurable (liveness on /status.php restarts healthy pods during dependency outages)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 72/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- helm, kubernetes
- Ambito
- devops
Direzione di ricerca
Inizia da templates/deployment.yaml e values.yaml: l’issue identifica tutti e sei i percorsi di probe hardcoded e le tre sezioni dei valori di probe. Aggiungi percorsi opzionali per ogni probe mantenendo /status.php come valore predefinito e aggiorna la README con i nuovi valori. Verifica che i rami del template per nginx e quelli senza nginx vengano renderizzati correttamente per tutte e tre le probe.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Description of the change
The chart renders all three probes (livenessProbe, readinessProbe, startupProbe) with a hardcoded httpGet.path: /status.php, while every other probe field is configurable:
templates/deployment.yaml— 6 hardcoded occurrences: lines 97 / 113 / 129 (non-nginx branch) and 158 / 174 / 190 (nginx.enabledbranch).values.yaml—livenessProbe/readinessProbe/startupProbeexpose onlyenabled,initialDelaySeconds,periodSeconds,timeoutSeconds,failureThreshold,successThreshold. There is nopath.
/status.php is served by PHP and depends on external systems (database, Redis, and the persistent volume mounted at /var/www/html). Using it for liveness means that when any dependency is slow or unavailable for longer than failureThreshold × periodSeconds (default 3 × 10s), the kubelet kills a perfectly healthy container. Kubernetes guidance is that liveness probes should not depend on external systems — dependency checks belong in readiness probes, where /status.php is exactly right and should stay.
What this caused in our cluster (chart 9.4.0, image.flavor: fpm + nginx.enabled: true, RKE2 v1.33)
| Time (UTC) | Event |
|---|---|
| 09:37:27 | the app persistent volume (RWX, CSI) became unreachable; every /status.php request exceeded timeoutSeconds: 5 |
| 09:37:57 | liveness failed 3× → kubelet sent SIGQUIT to the nginx sidecar (Container nextcloud-nginx failed liveness probe, will be restarted) |
| 09:37:57 → | nginx stopped accepting connections but never exited — its workers were blocked in uninterruptible I/O against the dead volume. nginx stayed running, restarts=0, startedAt unchanged, port 80 connection refused (×146 events) |
| +20 min | the pod could not be restarted by the kubelet and could not be deleted without --force --grace-period=0; the service stayed down until an operator intervened |
A readiness-only failure would have self-healed: as soon as the dependency returned, readiness would pass again and the pod would be re-added to the Service. With liveness on the same dependency-aware endpoint, the outage was converted into a wedged pod.
Mitigations available to operators today via values: none, other than livenessProbe.enabled: false (which also gives up restarting a genuinely deadlocked process). We verified that neither image.flavor nor the nginx.* values offer a probe path, and that the default files/nginx.config.tpl ships no dependency-free health endpoint.
Proposal 1 — minimal, backward compatible (requested)
Expose an optional path per probe and default it to the current value, so existing installations are unaffected:
# values.yaml
livenessProbe:
enabled: true
path: /status.php # NEW – optional, defaults to /status.php
initialDelaySeconds: 10
...
readinessProbe:
enabled: true
path: /status.php # NEW – optional
...
startupProbe:
enabled: false
path: /status.php # NEW – optional
...
--- a/charts/nextcloud/templates/deployment.yaml
+++ b/charts/nextcloud/templates/deployment.yaml
@@
{{- with .Values.livenessProbe }}
{{- if .enabled }}
livenessProbe:
httpGet:
- path: /status.php
+ path: {{ .path | default "/status.php" }}
port: {{ $.Values.nextcloud.containerPort }}
(same for readinessProbe / startupProbe, in both the non-nginx and the nginx.enabled branches.)
This alone lets operators implement the correct split — liveness = process liveness, readiness = dependencies — without patching the Deployment after every helm upgrade.
Proposal 2 — recommended follow-up: ship a dependency-free endpoint for liveness
Add a static location to the default nginx config (files/nginx.config.tpl, inside the existing server {} block):
location = /healthz {
access_log off;
return 200;
}
and default liveness to /healthz (readiness stays on /status.php). This endpoint is answered by nginx itself — no PHP, no database, no Redis, no persistent volume — so it fails only when the web server process is actually unable to serve.
Because this changes restart behaviour, it should land as an explicit, documented change (chart minor/major + README note), and it only applies where nginx is present:
nginx.enabled: true→ the probe targets the nginx container,/healthzis served there ✅nginx.enabled: false(mod_php / fpm without sidecar) → probes target the app container, so/healthzwould not exist. For these setups Proposal 1 (configurable path) is the useful part, or/healthzcould be provided as a static file / Apache snippet as a separate piece of work.
Workaround available today (worth documenting either way)
nginx:
config:
serverBlockCustom: |
client_max_body_size 512M;
client_body_timeout 300s;
fastcgi_buffers 64 4K;
fastcgi_read_timeout 3600s;
location = /healthz { access_log off; return 200; }
Note: serverBlockCustom is a full replacement of the default block (so the four default lines must be repeated), and it is injected into the server {} block of default.conf — a bare location therefore works there. The sibling key nginx.config.custom renders a separate file (zz-custom.conf) that is included in the http context, where a bare location would make nginx fail to start.
The endpoint exists afterwards, but making the kubelet use it still requires patching the Deployment, and that patch is reverted by the next Helm upgrade:
kubectl -n <ns> patch deploy nextcloud --type=strategic \
-p '{"spec":{"template":{"spec":{"containers":[{"name":"nextcloud-nginx","livenessProbe":{"httpGet":{"path":"/healthz"}}}]}}}}'
Benefits
- Removes a well-known Kubernetes anti-pattern (liveness depending on external systems) from the default configuration.
- Dependency outages (DB failover, Redis restart, NFS/RWX/CSI hiccup, storage maintenance) stop turning into container restarts and, in the worst case, wedged pods.
- Operators regain a real liveness signal: a deadlocked nginx/php-fpm is still detected and restarted, without coupling that decision to storage/database availability.
- Proposal 1 is fully backward compatible and useful on its own — it also helps with reverse proxies / other orchestrators that need a different probe path.
Possible drawbacks
- Proposal 2 changes default behaviour: pods are no longer restarted when
/status.phpfails. That is the intent, but it is a behavioural change and should be called out in the chart changelog and README. - A
/healthzanswered by nginx alone cannot detect a wedged php-fpm pool. Mitigation: keepstartupProbe/readinessProbeon/status.php(so traffic stops when PHP is unhealthy) — optionally recommendpm.max_requests/request_terminate_timeoutin the docs. - One more endpoint is exposed; it returns only
200and reveals nothing sensitive, but it is routable through the ingress unless restricted. - Chart surface grows by three optional values.
Additional information
Environment
- Kubernetes distribution: RKE2 v1.33 (guest cluster of a Harvester/HCI deployment)
- Helm: chart 9.4.0 (appVersion 35.0.1), managed through the Rancher UI
image.flavor: fpm+nginx.enabled: true;/var/www/html(code/config) on an RWX CSI volume, user data on a separate NFS CSI volume- Probes: chart defaults (
timeoutSeconds: 5,periodSeconds: 10,failureThreshold: 3)
Relevant chart facts (verified against main on 2026-10-06, chart 9.4.0 / appVersion 35.0.1)
templates/deployment.yamllines 97 / 113 / 129 / 158 / 174 / 190 →path: /status.php(hardcoded, both branches).values.yamllivenessProbe/readinessProbe/startupProbe→ nopathkey.files/nginx.config.tpl→ no health endpoint.nginx.config.serverBlockCustomis injected inside theserver {}block (files/nginx.config.tpl, line 53) — so it can carry alocation;nginx.config.customrenderszz-custom.confinto/etc/nginx/conf.d/(httpcontext), where a barelocationis invalid.- The shipped default
serverBlockCustomcontainsfastcgi_read_timeout 3600s;(server level), so a stuck PHP upstream can hold nginx workers for up to an hour — this is what made the graceful-shutdown hang in our incident possible. worker_shutdown_timeout(which would bound that shutdown) is amain-context directive and therefore cannot be injected throughnginx.config.customorserverBlockCustom; only a fullnginx.confreplacement would do it.
I'm happy to submit a PR for Proposal 1 (small and backward compatible: values.yaml + templates/deployment.yaml + one README row).
- Lingua principale
- Go Template
- Stelle
- 536
- Fork
- 316
- Merge medio
- 4g 16h
- PR unite (30g)
- 2
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di nextcloud/helm
-
No native support for REDIS_USER enviroment varForse già presa @jholmes802 l’ha presa 1 giorno fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Failed to inspect imageAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
nextcloud.openmetrics.allowedClients is silently ignored unless nextcloud.configs is setForse già presa @JanWelker l’ha presa 17 giorni fa. Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 78/100
-
External Redis not working: redis-session.ini: Permission deniedForse già presa @antoinetran l’ha presa 22 giorni fa. Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
Tutte le issue di nextcloud/helm
Issue simili
-
lens:agent lens:process process
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
thebristolsound/birdbrain#1771 ·
I maintainer di solito rispondono entro 1 giorno
-
ci-install-db-tools stall-case tests flake: stalled apt-get can be killed before it logs its callApertaeffort:low model:light plan planner:opus-5-5 tests
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
I maintainer di solito rispondono entro 1 giorno
-
component/test-automation work/tech-debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
bcgov/bc-wallet-mobile#4845 ·
I maintainer di solito rispondono entro 1 giorno
-
first
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
AcademySoftwareFoundation/rmtc#54 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
stac-utils/stac-fastapi-elasticsearch-opensearch#932 ·
I maintainer di solito rispondono entro 1 giorno