Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Make probe paths configurable (liveness on /status.php restarts healthy pods during dependency outages)

Aperta Adatta ai principianti
#890 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
72/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
helm, kubernetes
Ambito
devops

Direzione di ricerca

Inizia da templates/deployment.yaml e values.yaml: l’issue identifica tutti e sei i percorsi di probe hardcoded e le tre sezioni dei valori di probe. Aggiungi percorsi opzionali per ogni probe mantenendo /status.php come valore predefinito e aggiorna la README con i nuovi valori. Verifica che i rami del template per nginx e quelli senza nginx vengano renderizzati correttamente per tutte e tre le probe.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Description of the change

The chart renders all three probes (livenessProbe, readinessProbe, startupProbe) with a hardcoded httpGet.path: /status.php, while every other probe field is configurable:

  • templates/deployment.yaml — 6 hardcoded occurrences: lines 97 / 113 / 129 (non-nginx branch) and 158 / 174 / 190 (nginx.enabled branch).
  • values.yaml — livenessProbe / readinessProbe / startupProbe expose only enabled, initialDelaySeconds, periodSeconds, timeoutSeconds, failureThreshold, successThreshold. There is no path.

/status.php is served by PHP and depends on external systems (database, Redis, and the persistent volume mounted at /var/www/html). Using it for liveness means that when any dependency is slow or unavailable for longer than failureThreshold × periodSeconds (default 3 × 10s), the kubelet kills a perfectly healthy container. Kubernetes guidance is that liveness probes should not depend on external systems — dependency checks belong in readiness probes, where /status.php is exactly right and should stay.

What this caused in our cluster (chart 9.4.0, image.flavor: fpm + nginx.enabled: true, RKE2 v1.33)
Time (UTC) Event
09:37:27 the app persistent volume (RWX, CSI) became unreachable; every /status.php request exceeded timeoutSeconds: 5
09:37:57 liveness failed 3× → kubelet sent SIGQUIT to the nginx sidecar (Container nextcloud-nginx failed liveness probe, will be restarted)
09:37:57 → nginx stopped accepting connections but never exited — its workers were blocked in uninterruptible I/O against the dead volume. nginx stayed running, restarts=0, startedAt unchanged, port 80 connection refused (×146 events)
+20 min the pod could not be restarted by the kubelet and could not be deleted without --force --grace-period=0; the service stayed down until an operator intervened

A readiness-only failure would have self-healed: as soon as the dependency returned, readiness would pass again and the pod would be re-added to the Service. With liveness on the same dependency-aware endpoint, the outage was converted into a wedged pod.

Mitigations available to operators today via values: none, other than livenessProbe.enabled: false (which also gives up restarting a genuinely deadlocked process). We verified that neither image.flavor nor the nginx.* values offer a probe path, and that the default files/nginx.config.tpl ships no dependency-free health endpoint.

Proposal 1 — minimal, backward compatible (requested)

Expose an optional path per probe and default it to the current value, so existing installations are unaffected:

# values.yaml
livenessProbe:
  enabled: true
  path: /status.php        # NEW – optional, defaults to /status.php
  initialDelaySeconds: 10
  ...
readinessProbe:
  enabled: true
  path: /status.php        # NEW – optional
  ...
startupProbe:
  enabled: false
  path: /status.php        # NEW – optional
  ...
--- a/charts/nextcloud/templates/deployment.yaml
+++ b/charts/nextcloud/templates/deployment.yaml
@@
           {{- with .Values.livenessProbe }}
           {{- if .enabled }}
           livenessProbe:
             httpGet:
-              path: /status.php
+              path: {{ .path | default "/status.php" }}
               port:  {{ $.Values.nextcloud.containerPort }}

(same for readinessProbe / startupProbe, in both the non-nginx and the nginx.enabled branches.)

This alone lets operators implement the correct split — liveness = process liveness, readiness = dependencies — without patching the Deployment after every helm upgrade.

Proposal 2 — recommended follow-up: ship a dependency-free endpoint for liveness

Add a static location to the default nginx config (files/nginx.config.tpl, inside the existing server {} block):

location = /healthz {
    access_log off;
    return 200;
}

and default liveness to /healthz (readiness stays on /status.php). This endpoint is answered by nginx itself — no PHP, no database, no Redis, no persistent volume — so it fails only when the web server process is actually unable to serve.

Because this changes restart behaviour, it should land as an explicit, documented change (chart minor/major + README note), and it only applies where nginx is present:

  • nginx.enabled: true → the probe targets the nginx container, /healthz is served there ✅
  • nginx.enabled: false (mod_php / fpm without sidecar) → probes target the app container, so /healthz would not exist. For these setups Proposal 1 (configurable path) is the useful part, or /healthz could be provided as a static file / Apache snippet as a separate piece of work.
Workaround available today (worth documenting either way)
nginx:
  config:
    serverBlockCustom: |
      client_max_body_size 512M;
      client_body_timeout 300s;
      fastcgi_buffers 64 4K;
      fastcgi_read_timeout 3600s;
      location = /healthz { access_log off; return 200; }

Note: serverBlockCustom is a full replacement of the default block (so the four default lines must be repeated), and it is injected into the server {} block of default.conf — a bare location therefore works there. The sibling key nginx.config.custom renders a separate file (zz-custom.conf) that is included in the http context, where a bare location would make nginx fail to start.

The endpoint exists afterwards, but making the kubelet use it still requires patching the Deployment, and that patch is reverted by the next Helm upgrade:

kubectl -n <ns> patch deploy nextcloud --type=strategic \
  -p '{"spec":{"template":{"spec":{"containers":[{"name":"nextcloud-nginx","livenessProbe":{"httpGet":{"path":"/healthz"}}}]}}}}'

Benefits

  • Removes a well-known Kubernetes anti-pattern (liveness depending on external systems) from the default configuration.
  • Dependency outages (DB failover, Redis restart, NFS/RWX/CSI hiccup, storage maintenance) stop turning into container restarts and, in the worst case, wedged pods.
  • Operators regain a real liveness signal: a deadlocked nginx/php-fpm is still detected and restarted, without coupling that decision to storage/database availability.
  • Proposal 1 is fully backward compatible and useful on its own — it also helps with reverse proxies / other orchestrators that need a different probe path.

Possible drawbacks

  • Proposal 2 changes default behaviour: pods are no longer restarted when /status.php fails. That is the intent, but it is a behavioural change and should be called out in the chart changelog and README.
  • A /healthz answered by nginx alone cannot detect a wedged php-fpm pool. Mitigation: keep startupProbe/readinessProbe on /status.php (so traffic stops when PHP is unhealthy) — optionally recommend pm.max_requests / request_terminate_timeout in the docs.
  • One more endpoint is exposed; it returns only 200 and reveals nothing sensitive, but it is routable through the ingress unless restricted.
  • Chart surface grows by three optional values.

Additional information

Environment

  • Kubernetes distribution: RKE2 v1.33 (guest cluster of a Harvester/HCI deployment)
  • Helm: chart 9.4.0 (appVersion 35.0.1), managed through the Rancher UI
  • image.flavor: fpm + nginx.enabled: true; /var/www/html (code/config) on an RWX CSI volume, user data on a separate NFS CSI volume
  • Probes: chart defaults (timeoutSeconds: 5, periodSeconds: 10, failureThreshold: 3)

Relevant chart facts (verified against main on 2026-10-06, chart 9.4.0 / appVersion 35.0.1)

  1. templates/deployment.yaml lines 97 / 113 / 129 / 158 / 174 / 190 → path: /status.php (hardcoded, both branches).
  2. values.yaml livenessProbe / readinessProbe / startupProbe → no path key.
  3. files/nginx.config.tpl → no health endpoint.
  4. nginx.config.serverBlockCustom is injected inside the server {} block (files/nginx.config.tpl, line 53) — so it can carry a location; nginx.config.custom renders zz-custom.conf into /etc/nginx/conf.d/ (http context), where a bare location is invalid.
  5. The shipped default serverBlockCustom contains fastcgi_read_timeout 3600s; (server level), so a stuck PHP upstream can hold nginx workers for up to an hour — this is what made the graceful-shutdown hang in our incident possible.
  6. worker_shutdown_timeout (which would bound that shutdown) is a main-context directive and therefore cannot be injected through nginx.config.custom or serverBlockCustom; only a full nginx.conf replacement would do it.

I'm happy to submit a PR for Proposal 1 (small and backward compatible: values.yaml + templates/deployment.yaml + one README row).

Lingua principale
Go Template
Stelle
536
Fork
316
Merge medio
4g 16h
PR unite (30g)
2

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di nextcloud/helm

Tutte le issue di nextcloud/helm

Issue simili

Altre issue su DevOps

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.