postgres in container with PID 1 aggregating orphaned processes leading to restarts & recovery
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- docker, kubernetes, postgres, shell
- Domain
- databases, devops, infrastructure
Research direction
Start with the reported Dockerfile example, the docker-entrypoint.sh command, and the linked dumb-init article. The issue does not name a repository file or test, so first confirm whether the requested scope is an image entrypoint change; done should include a repository-supported remediation and validation that the probe-related restarts no longer occur.
Written by the indexing model from the issue text.
Description
The postgres in docker hosted on K8S infrastructure with periodic health and rediness checks experiences regular restarts of the database.
Logs show 'postgres "server process <PID>exited with exit code 2" ' followed by restart and recovery of the database.
happening with a variety of stable postgres versions (12, 14, 16) on the same system
deployed using kubernetes, we monitor health with exec probe
livenessProbe:
exec:
command:
- pg_isready
- -U
- dmp-admin
- -d
- dmp-entity
timeoutSeconds: 1
as well as readiness probe
readinessProbe:
exec:
command:
- /bin/bash
- -c
- pg_isready -U dmp-admin -d dmp-entity && [ ! -f /var/lib/postgresql/backup/pgdump_backup.velero.sql ]
timeoutSeconds: 1
Investigation shows that the postgres db is fine - no exit 2 occurences.
Instead a regular health monitoring process that includes a /bin/bash -c pg_isready .. causes the problem.
Root cause is described here: https://www.cybertec-postgresql.com/en/docker-sudden-death-for-postgresql/ (thanks to laurentz albe)
The postgres in docker setup runs the default process (postgres) as root process of the container with PID 1.
The main postgres container is a health manager/monitor for all other spawned worker processes.
Sudden exits of worker processes lead to DB restart - remediating possible shared memory corruption.
However, orphaned other processes will get the root process as parent process (PID 1) being our postgres main entrypoint.
These processes get orphaned due to to timeout of the monitoring environment.
pg_isready is known to return exit code 2 when not able to connect.
Suggested remediation is in the referenced article: start the container using dum-init.
Example patch we deploy to remediate consists of installing the dumb-init package (using apt) and extending the entrypoint to use dumb-init as the main process (PID 1)
FROM postgres:12
RUN apt update && apt install -y dumb-init && apt clean
ENTRYPOINT ["/usr/bin/dumb-init", "docker-entrypoint.sh"]
CMD ["postgres"]
- Dominant language
- Shell
- Stars
- 2.5k
- Forks
- 1.2k
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from docker-library/postgres
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
docker-library/postgres#1420 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
docker-library/postgres#1419 · 2 comments ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
docker-library/postgres#1389 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
docker-library/postgres#1356 · 5 comments · 7 reactions ·
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
docker-library/postgres#1355 · 10 comments · 11 reactions ·
All issues in docker-library/postgres
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
obra/superpowers#2445 ·
Maintainers usually reply within 5 days
-
package-update
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
oSoWoSo/vOid_Community_repOsitory#240 · 1 comment ·
Maintainers usually reply within 1 day
-
ai-inspected
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
[Bug]: atuin doctor reports "Hub (authenticated)" while the client syncs with a self-hosted serverOpen
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day