Make probe paths configurable (liveness on /status.php restarts healthy pods during dependency outages)
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 72/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- helm, kubernetes
- Domain
- devops
Research direction
Start with templates/deployment.yaml and values.yaml: the issue identifies all six hardcoded probe paths and the three probe value sections. Add optional per-probe paths while preserving /status.php as the default, and update the README with the new values. Verify both nginx and non-nginx template branches render correctly for all three probes.
Written by the indexing model from the issue text.
Description
Description of the change
The chart renders all three probes (livenessProbe, readinessProbe, startupProbe) with a hardcoded httpGet.path: /status.php, while every other probe field is configurable:
templates/deployment.yaml— 6 hardcoded occurrences: lines 97 / 113 / 129 (non-nginx branch) and 158 / 174 / 190 (nginx.enabledbranch).values.yaml—livenessProbe/readinessProbe/startupProbeexpose onlyenabled,initialDelaySeconds,periodSeconds,timeoutSeconds,failureThreshold,successThreshold. There is nopath.
/status.php is served by PHP and depends on external systems (database, Redis, and the persistent volume mounted at /var/www/html). Using it for liveness means that when any dependency is slow or unavailable for longer than failureThreshold × periodSeconds (default 3 × 10s), the kubelet kills a perfectly healthy container. Kubernetes guidance is that liveness probes should not depend on external systems — dependency checks belong in readiness probes, where /status.php is exactly right and should stay.
What this caused in our cluster (chart 9.4.0, image.flavor: fpm + nginx.enabled: true, RKE2 v1.33)
| Time (UTC) | Event |
|---|---|
| 09:37:27 | the app persistent volume (RWX, CSI) became unreachable; every /status.php request exceeded timeoutSeconds: 5 |
| 09:37:57 | liveness failed 3× → kubelet sent SIGQUIT to the nginx sidecar (Container nextcloud-nginx failed liveness probe, will be restarted) |
| 09:37:57 → | nginx stopped accepting connections but never exited — its workers were blocked in uninterruptible I/O against the dead volume. nginx stayed running, restarts=0, startedAt unchanged, port 80 connection refused (×146 events) |
| +20 min | the pod could not be restarted by the kubelet and could not be deleted without --force --grace-period=0; the service stayed down until an operator intervened |
A readiness-only failure would have self-healed: as soon as the dependency returned, readiness would pass again and the pod would be re-added to the Service. With liveness on the same dependency-aware endpoint, the outage was converted into a wedged pod.
Mitigations available to operators today via values: none, other than livenessProbe.enabled: false (which also gives up restarting a genuinely deadlocked process). We verified that neither image.flavor nor the nginx.* values offer a probe path, and that the default files/nginx.config.tpl ships no dependency-free health endpoint.
Proposal 1 — minimal, backward compatible (requested)
Expose an optional path per probe and default it to the current value, so existing installations are unaffected:
# values.yaml
livenessProbe:
enabled: true
path: /status.php # NEW – optional, defaults to /status.php
initialDelaySeconds: 10
...
readinessProbe:
enabled: true
path: /status.php # NEW – optional
...
startupProbe:
enabled: false
path: /status.php # NEW – optional
...
--- a/charts/nextcloud/templates/deployment.yaml
+++ b/charts/nextcloud/templates/deployment.yaml
@@
{{- with .Values.livenessProbe }}
{{- if .enabled }}
livenessProbe:
httpGet:
- path: /status.php
+ path: {{ .path | default "/status.php" }}
port: {{ $.Values.nextcloud.containerPort }}
(same for readinessProbe / startupProbe, in both the non-nginx and the nginx.enabled branches.)
This alone lets operators implement the correct split — liveness = process liveness, readiness = dependencies — without patching the Deployment after every helm upgrade.
Proposal 2 — recommended follow-up: ship a dependency-free endpoint for liveness
Add a static location to the default nginx config (files/nginx.config.tpl, inside the existing server {} block):
location = /healthz {
access_log off;
return 200;
}
and default liveness to /healthz (readiness stays on /status.php). This endpoint is answered by nginx itself — no PHP, no database, no Redis, no persistent volume — so it fails only when the web server process is actually unable to serve.
Because this changes restart behaviour, it should land as an explicit, documented change (chart minor/major + README note), and it only applies where nginx is present:
nginx.enabled: true→ the probe targets the nginx container,/healthzis served there ✅nginx.enabled: false(mod_php / fpm without sidecar) → probes target the app container, so/healthzwould not exist. For these setups Proposal 1 (configurable path) is the useful part, or/healthzcould be provided as a static file / Apache snippet as a separate piece of work.
Workaround available today (worth documenting either way)
nginx:
config:
serverBlockCustom: |
client_max_body_size 512M;
client_body_timeout 300s;
fastcgi_buffers 64 4K;
fastcgi_read_timeout 3600s;
location = /healthz { access_log off; return 200; }
Note: serverBlockCustom is a full replacement of the default block (so the four default lines must be repeated), and it is injected into the server {} block of default.conf — a bare location therefore works there. The sibling key nginx.config.custom renders a separate file (zz-custom.conf) that is included in the http context, where a bare location would make nginx fail to start.
The endpoint exists afterwards, but making the kubelet use it still requires patching the Deployment, and that patch is reverted by the next Helm upgrade:
kubectl -n <ns> patch deploy nextcloud --type=strategic \
-p '{"spec":{"template":{"spec":{"containers":[{"name":"nextcloud-nginx","livenessProbe":{"httpGet":{"path":"/healthz"}}}]}}}}'
Benefits
- Removes a well-known Kubernetes anti-pattern (liveness depending on external systems) from the default configuration.
- Dependency outages (DB failover, Redis restart, NFS/RWX/CSI hiccup, storage maintenance) stop turning into container restarts and, in the worst case, wedged pods.
- Operators regain a real liveness signal: a deadlocked nginx/php-fpm is still detected and restarted, without coupling that decision to storage/database availability.
- Proposal 1 is fully backward compatible and useful on its own — it also helps with reverse proxies / other orchestrators that need a different probe path.
Possible drawbacks
- Proposal 2 changes default behaviour: pods are no longer restarted when
/status.phpfails. That is the intent, but it is a behavioural change and should be called out in the chart changelog and README. - A
/healthzanswered by nginx alone cannot detect a wedged php-fpm pool. Mitigation: keepstartupProbe/readinessProbeon/status.php(so traffic stops when PHP is unhealthy) — optionally recommendpm.max_requests/request_terminate_timeoutin the docs. - One more endpoint is exposed; it returns only
200and reveals nothing sensitive, but it is routable through the ingress unless restricted. - Chart surface grows by three optional values.
Additional information
Environment
- Kubernetes distribution: RKE2 v1.33 (guest cluster of a Harvester/HCI deployment)
- Helm: chart 9.4.0 (appVersion 35.0.1), managed through the Rancher UI
image.flavor: fpm+nginx.enabled: true;/var/www/html(code/config) on an RWX CSI volume, user data on a separate NFS CSI volume- Probes: chart defaults (
timeoutSeconds: 5,periodSeconds: 10,failureThreshold: 3)
Relevant chart facts (verified against main on 2026-10-06, chart 9.4.0 / appVersion 35.0.1)
templates/deployment.yamllines 97 / 113 / 129 / 158 / 174 / 190 →path: /status.php(hardcoded, both branches).values.yamllivenessProbe/readinessProbe/startupProbe→ nopathkey.files/nginx.config.tpl→ no health endpoint.nginx.config.serverBlockCustomis injected inside theserver {}block (files/nginx.config.tpl, line 53) — so it can carry alocation;nginx.config.customrenderszz-custom.confinto/etc/nginx/conf.d/(httpcontext), where a barelocationis invalid.- The shipped default
serverBlockCustomcontainsfastcgi_read_timeout 3600s;(server level), so a stuck PHP upstream can hold nginx workers for up to an hour — this is what made the graceful-shutdown hang in our incident possible. worker_shutdown_timeout(which would bound that shutdown) is amain-context directive and therefore cannot be injected throughnginx.config.customorserverBlockCustom; only a fullnginx.confreplacement would do it.
I'm happy to submit a PR for Proposal 1 (small and backward compatible: values.yaml + templates/deployment.yaml + one README row).
- Dominant language
- Go Template
- Stars
- 536
- Forks
- 316
- Avg merge
- 4d 16h
- Merged PRs (30d)
- 2
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from nextcloud/helm
-
No native support for REDIS_USER enviroment varPossibly taken @jholmes802 claimed this 3 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
nextcloud.openmetrics.allowedClients is silently ignored unless nextcloud.configs is setPossibly taken @JanWelker claimed this 20 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 78/100
-
External Redis not working: redis-session.ini: Permission deniedPossibly taken @antoinetran claimed this 24 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 55/100
-
Difficulty 3/5 1-2 days Newbie friendliness 55/100
Similar issues
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOpenneeds-triage
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
krkn-chaos/krkn#1627 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
int128/typescript-action#1575 ·
Maintainers usually reply within 1 day
-
agent/sec-check hive/hosted-available-lke648397-260827-5n31 security
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
femiwiki/docker-mediawiki#1497 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
EPFL-ENAC/co2-calculator#3055 ·
Maintainers usually reply within 1 day