Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Make probe paths configurable (liveness on /status.php restarts healthy pods during dependency outages)

Open Beginner friendly
#890 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
72/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
helm, kubernetes
Domain
devops

Research direction

Start with templates/deployment.yaml and values.yaml: the issue identifies all six hardcoded probe paths and the three probe value sections. Add optional per-probe paths while preserving /status.php as the default, and update the README with the new values. Verify both nginx and non-nginx template branches render correctly for all three probes.

Written by the indexing model from the issue text.

Description

Description of the change

The chart renders all three probes (livenessProbe, readinessProbe, startupProbe) with a hardcoded httpGet.path: /status.php, while every other probe field is configurable:

  • templates/deployment.yaml — 6 hardcoded occurrences: lines 97 / 113 / 129 (non-nginx branch) and 158 / 174 / 190 (nginx.enabled branch).
  • values.yaml — livenessProbe / readinessProbe / startupProbe expose only enabled, initialDelaySeconds, periodSeconds, timeoutSeconds, failureThreshold, successThreshold. There is no path.

/status.php is served by PHP and depends on external systems (database, Redis, and the persistent volume mounted at /var/www/html). Using it for liveness means that when any dependency is slow or unavailable for longer than failureThreshold × periodSeconds (default 3 × 10s), the kubelet kills a perfectly healthy container. Kubernetes guidance is that liveness probes should not depend on external systems — dependency checks belong in readiness probes, where /status.php is exactly right and should stay.

What this caused in our cluster (chart 9.4.0, image.flavor: fpm + nginx.enabled: true, RKE2 v1.33)
Time (UTC) Event
09:37:27 the app persistent volume (RWX, CSI) became unreachable; every /status.php request exceeded timeoutSeconds: 5
09:37:57 liveness failed 3× → kubelet sent SIGQUIT to the nginx sidecar (Container nextcloud-nginx failed liveness probe, will be restarted)
09:37:57 → nginx stopped accepting connections but never exited — its workers were blocked in uninterruptible I/O against the dead volume. nginx stayed running, restarts=0, startedAt unchanged, port 80 connection refused (×146 events)
+20 min the pod could not be restarted by the kubelet and could not be deleted without --force --grace-period=0; the service stayed down until an operator intervened

A readiness-only failure would have self-healed: as soon as the dependency returned, readiness would pass again and the pod would be re-added to the Service. With liveness on the same dependency-aware endpoint, the outage was converted into a wedged pod.

Mitigations available to operators today via values: none, other than livenessProbe.enabled: false (which also gives up restarting a genuinely deadlocked process). We verified that neither image.flavor nor the nginx.* values offer a probe path, and that the default files/nginx.config.tpl ships no dependency-free health endpoint.

Proposal 1 — minimal, backward compatible (requested)

Expose an optional path per probe and default it to the current value, so existing installations are unaffected:

# values.yaml
livenessProbe:
  enabled: true
  path: /status.php        # NEW – optional, defaults to /status.php
  initialDelaySeconds: 10
  ...
readinessProbe:
  enabled: true
  path: /status.php        # NEW – optional
  ...
startupProbe:
  enabled: false
  path: /status.php        # NEW – optional
  ...
--- a/charts/nextcloud/templates/deployment.yaml
+++ b/charts/nextcloud/templates/deployment.yaml
@@
           {{- with .Values.livenessProbe }}
           {{- if .enabled }}
           livenessProbe:
             httpGet:
-              path: /status.php
+              path: {{ .path | default "/status.php" }}
               port:  {{ $.Values.nextcloud.containerPort }}

(same for readinessProbe / startupProbe, in both the non-nginx and the nginx.enabled branches.)

This alone lets operators implement the correct split — liveness = process liveness, readiness = dependencies — without patching the Deployment after every helm upgrade.

Proposal 2 — recommended follow-up: ship a dependency-free endpoint for liveness

Add a static location to the default nginx config (files/nginx.config.tpl, inside the existing server {} block):

location = /healthz {
    access_log off;
    return 200;
}

and default liveness to /healthz (readiness stays on /status.php). This endpoint is answered by nginx itself — no PHP, no database, no Redis, no persistent volume — so it fails only when the web server process is actually unable to serve.

Because this changes restart behaviour, it should land as an explicit, documented change (chart minor/major + README note), and it only applies where nginx is present:

  • nginx.enabled: true → the probe targets the nginx container, /healthz is served there ✅
  • nginx.enabled: false (mod_php / fpm without sidecar) → probes target the app container, so /healthz would not exist. For these setups Proposal 1 (configurable path) is the useful part, or /healthz could be provided as a static file / Apache snippet as a separate piece of work.
Workaround available today (worth documenting either way)
nginx:
  config:
    serverBlockCustom: |
      client_max_body_size 512M;
      client_body_timeout 300s;
      fastcgi_buffers 64 4K;
      fastcgi_read_timeout 3600s;
      location = /healthz { access_log off; return 200; }

Note: serverBlockCustom is a full replacement of the default block (so the four default lines must be repeated), and it is injected into the server {} block of default.conf — a bare location therefore works there. The sibling key nginx.config.custom renders a separate file (zz-custom.conf) that is included in the http context, where a bare location would make nginx fail to start.

The endpoint exists afterwards, but making the kubelet use it still requires patching the Deployment, and that patch is reverted by the next Helm upgrade:

kubectl -n <ns> patch deploy nextcloud --type=strategic \
  -p '{"spec":{"template":{"spec":{"containers":[{"name":"nextcloud-nginx","livenessProbe":{"httpGet":{"path":"/healthz"}}}]}}}}'

Benefits

  • Removes a well-known Kubernetes anti-pattern (liveness depending on external systems) from the default configuration.
  • Dependency outages (DB failover, Redis restart, NFS/RWX/CSI hiccup, storage maintenance) stop turning into container restarts and, in the worst case, wedged pods.
  • Operators regain a real liveness signal: a deadlocked nginx/php-fpm is still detected and restarted, without coupling that decision to storage/database availability.
  • Proposal 1 is fully backward compatible and useful on its own — it also helps with reverse proxies / other orchestrators that need a different probe path.

Possible drawbacks

  • Proposal 2 changes default behaviour: pods are no longer restarted when /status.php fails. That is the intent, but it is a behavioural change and should be called out in the chart changelog and README.
  • A /healthz answered by nginx alone cannot detect a wedged php-fpm pool. Mitigation: keep startupProbe/readinessProbe on /status.php (so traffic stops when PHP is unhealthy) — optionally recommend pm.max_requests / request_terminate_timeout in the docs.
  • One more endpoint is exposed; it returns only 200 and reveals nothing sensitive, but it is routable through the ingress unless restricted.
  • Chart surface grows by three optional values.

Additional information

Environment

  • Kubernetes distribution: RKE2 v1.33 (guest cluster of a Harvester/HCI deployment)
  • Helm: chart 9.4.0 (appVersion 35.0.1), managed through the Rancher UI
  • image.flavor: fpm + nginx.enabled: true; /var/www/html (code/config) on an RWX CSI volume, user data on a separate NFS CSI volume
  • Probes: chart defaults (timeoutSeconds: 5, periodSeconds: 10, failureThreshold: 3)

Relevant chart facts (verified against main on 2026-10-06, chart 9.4.0 / appVersion 35.0.1)

  1. templates/deployment.yaml lines 97 / 113 / 129 / 158 / 174 / 190 → path: /status.php (hardcoded, both branches).
  2. values.yaml livenessProbe / readinessProbe / startupProbe → no path key.
  3. files/nginx.config.tpl → no health endpoint.
  4. nginx.config.serverBlockCustom is injected inside the server {} block (files/nginx.config.tpl, line 53) — so it can carry a location; nginx.config.custom renders zz-custom.conf into /etc/nginx/conf.d/ (http context), where a bare location is invalid.
  5. The shipped default serverBlockCustom contains fastcgi_read_timeout 3600s; (server level), so a stuck PHP upstream can hold nginx workers for up to an hour — this is what made the graceful-shutdown hang in our incident possible.
  6. worker_shutdown_timeout (which would bound that shutdown) is a main-context directive and therefore cannot be injected through nginx.config.custom or serverBlockCustom; only a full nginx.conf replacement would do it.

I'm happy to submit a PR for Proposal 1 (small and backward compatible: values.yaml + templates/deployment.yaml + one README row).

Dominant language
Go Template
Stars
536
Forks
316
Avg merge
4d 16h
Merged PRs (30d)
2

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from nextcloud/helm

All issues in nextcloud/helm

Similar issues

More DevOps issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.