Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

`supabase/postgres:17.6.1.105` (amd64) — backend SIGSEGV during a heavy pgTAP file, CI-CPU-specific

オープン
#2,467 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
30/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
活発
技術スタック
docker, github-actions, postgresql

調査の方向性

まず、指定された PGOPTIONS 設定を使用し、AVX-512 対応の amd64 GitHub Actions runner 上で、完全に自己完結した pgTAP ファイルを public.ecr.aws/supabase/postgres:17.6.1.105 に対して実行します。runner の lscpu データと backend core または backtrace を収集し、その後、失敗がイメージのビルドまたは dispatch path に起因するものかどうかを判断します。完了とは、障害が確認されるか、文書化された緩和策を伴う再現可能な縮小ケースが得られることを意味します。

索引モデルが issue の本文から書いたものです。

説明

Summary

The amd64 public.ecr.aws/supabase/postgres:17.6.1.105 image deterministically SIGSEGVs a backend process (signal 11) while running one large single-transaction pgTAP test file — but only on GitHub Actions ubuntu-latest runners. The byte-identical image on a local amd64 workstation (Intel i7-8850H, AVX2) runs the same file cleanly every time, and our aarch64 production build is unaffected. This points to a CPU-microarchitecture-dependent code path in the amd64 17.6.1.105 build (the runner CPU is AVX-512-capable; our local CPU is not — AVX-512 is our leading hypothesis but is not core-dump-confirmed).

Environment

Image / build CPU Result
GitHub Actions ubuntu-latest public.ecr.aws/supabase/postgres:17.6.1.105 (amd64) AVX-512-capable x86_64 (Xeon/EPYC-class) SIGSEGV, 6/6 runs
Local workstation same image (amd64) Intel i7-8850H (AVX2, no AVX-512) clean, every run
Supabase production aarch64 build — PostgreSQL 17.6 on aarch64-unknown-linux-gnu ARM64 clean
  • Local version(): PostgreSQL 17.6 on x86_64-pc-linux-gnu, compiled by gcc (GCC) 15.2.0, 64-bit
  • Invocation: docker exec -i supabase_db_<project> psql -U postgres -d postgres -v ON_ERROR_STOP=1 < <file>.sql (Supabase CLI local stack; the file runs inside one BEGIN … ROLLBACK).

Observed behavior

A single pgTAP file (plan(363), ~4300 lines, one long transaction in one backend) crashes the backend mid-run:

LOG:  server process (PID …) was terminated by signal 11: Segmentation fault
LOG:  terminating any other active server connections
… database system is in recovery mode

The cluster then enters recovery and takes down the sibling test files running after it.

What it is NOT (ruled out with evidence)

  • Not OOM — signal 11 (not 9); dmesg is empty; adding swap did not help.
  • Not JIT — jit=off set per-connection via PGOPTIONS (verified active via SHOW jit); still crashed.
  • Not the outer plan shape — forcing plans with GUCs (enable_bitmapscan/hashagg/hashjoin=off, max_parallel_workers_per_gather=0) did not change the crash.
  • Not the SQL of the crashing statement — the crash lands on a statement (SELECT * FROM a SECURITY DEFINER SETOF-uuid helper) invoked by a role whose EXECUTE on that function is revoked, so it errors 42501 at the ACL check before the function body runs. The identical call by a privileged role earlier in the same file (and the same query shape ~17× before) runs clean. So the fault is not inside that function's execution — it is a delayed manifestation of state accumulated earlier in the backend.

The accumulation that appears to trigger it

The file, in one long-lived backend, builds up an unusually heavy load before the crash: hundreds of set_eq / EXCEPT-based pgTAP assertions (each materializing temp result sets), a BYPASSRLS role, and repeated trigger DDL / RLS-policy evaluation, all in one transaction. The corruption seems laid down by that accumulation and surfaces one statement later. On a fresh backend (e.g. splitting the file so the tail runs in its own psql/backend) the crash locus changes — consistent with private-backend-memory corruption rather than a specific query.

Reproduction

The full test file is self-contained pgTAP and reproduces 6/6 on an AVX-512-capable amd64 runner. We do not yet have a reduced repro (bisecting requires iterating on the AVX-512 hardware, which we only have via CI). We can share the full file privately, or work with you on a minimal case. To attempt repro:

  1. supabase start (or run public.ecr.aws/supabase/postgres:17.6.1.105 directly) on an AVX-512-capable amd64 host.
  2. Apply a schema with RLS + a SECURITY DEFINER set-returning helper, then run a long single-transaction pgTAP file that accumulates several hundred set_eq/EXCEPT assertions + a BYPASSRLS role + trigger DDL.
  3. Observe the backend SIGSEGV partway through.

A build compiled -O2 with CPU-dispatched SIMD (or a gcc 15.2.0 codegen issue on AVX-512) is the working hypothesis; a backtrace/core from the crashing backend on the runner would confirm the faulting routine.

Impact & current workaround

  • We have quarantined the one file from CI and gated it locally on AVX2; all other tests pass. No production impact — production is aarch64 (a different build with no x86 SIMD path), and the helper executes correctly there.
  • Asking here because it is your image build: a fix in the amd64 17.6.1.x build would let everyone re-enable such tests on standard CI runners.

Questions

  1. Is a SIGSEGV in the 17.6.1.105 amd64 build on AVX-512 hardware a known issue?
  2. Any recommended mitigation short of pinning an older/newer 17.x image (a build flag, a disabled dispatch path)?
  3. Would a full repro file + a runner lscpu be useful, or do you want a reduced case first?
主要言語
Nix
スター
1.8k
フォーク
268
平均マージ
1日 20時間
マージ済み PR(30日)
32

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

supabase/postgres のほかの issue

supabase/postgres の issue をすべて見る

似ている issue

DevOps の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。