Start CDC replay before deferred index creation
Nessuno ha ancora preso questa issue.
- #11 di @tbarbugli — chiusa senza merge
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- go
- Ambito
- databases, distributed-systems
Direzione di ricerca
Inizia tracciando i percorsi esistenti per COPY, il replay di CDC, la finalizzazione di indici e vincoli e lo stato di ripresa persistente; usa PR #9 come punto di ingresso di riferimento per le ricevute persistenti fuori ordine. Il completamento richiede pipeline coordinate di replay e di indicizzazione differita, un recupero sicuro dopo un crash, il mantenimento dell'ordine delle transazioni e una verifica end-to-end che i dati e gli indici del cutover corrispondano.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
pgmigrate currently captures CDC throughout schema restore, COPY, index/constraint creation, and target vacuuming, but does not start applying it until all of that work finishes. On large migrations, secondary index builds can therefore produce a very large CDC backlog and a long catch-up phase even though the copied tables are already usable for replay.
Proposed flow
capture CDC ─────────────────────────────────────────────>
pre-data → COPY → replay-critical indexes → CDC replay ──>
└→ deferred indexes concurrently
The first iteration should start replay after all tables finish COPY, but before secondary index and deferred schema work completes.
1. Classify indexes before COPY
Inventory and persist index definitions early, dividing them into:
- Replay-critical: primary-key and replica-identity indexes required for efficient
UPDATE/DELETEreplay. - Deferred: secondary indexes, non-identity unique indexes, and other post-data objects.
Partitioned indexes need leaf indexes built and validated before their parent index is attached.
2. Establish a replay-ready barrier
After all COPY transactions commit:
- Build replay-critical indexes using normal
CREATE INDEXwhile replay is still stopped. - Attach required primary-key and replica-identity constraints.
- Run
ANALYZEon copied tables. - Persist a replay-ready marker.
Replay must not begin without the key indexes, otherwise indexed updates and deletes can degrade into table scans.
3. Run replay and finalization concurrently
After the replay-ready barrier, supervise two pipelines together:
- CDC replay/catch-up.
- Deferred schema and index finalization.
A failure in either pipeline cancels the other. Reaching the initial CDC boundary is not sufficient to enter follow; required schema finalization must also be complete.
4. Make deferred index creation replay-safe
Once replay is active:
- Use
CREATE INDEX CONCURRENTLYfor deferred indexes. - Use a small, independent index worker pool.
- Persist
pending,building,valid, andcompletedstate. - On restart, identify owned invalid indexes and safely remove/retry them.
- Never run an ordinary long
CREATE INDEXagainst a table receiving replay.
5. Handle constraints deliberately
- Primary-key and replica-identity constraints must exist before replay.
- Foreign keys can be restored as
NOT VALID. - Do not validate foreign keys while out-of-order replay may expose temporary referential inconsistencies.
- Validate at a canonical replay boundary or during final drain.
- Preserve the existing replica-role behavior for triggers and rules.
6. Prioritize replay over background work
- Reserve target connections for replay.
- Default deferred index concurrency to one.
- Stop admitting new index builds when CDC lag exceeds a threshold.
- Resume index work when lag recovers.
- Treat vacuum and additional analysis as lower-priority background work.
7. Persist independent resume state
Track these milestones independently:
- COPY complete.
- Replay-critical indexes complete.
- Replay started.
- Deferred indexes complete.
- Constraints validated.
- Initial CDC boundary reached.
Resume must reconstruct both concurrent pipelines from durable state. follow and cutover remain gated on all required milestones.
Follow-up: per-table COPY handoff
If COPY itself remains the dominant source of backlog, allow each table to become replay-ready independently:
table A: COPY → key index → replay-ready
table B: COPY ─────────→ key index → replay-ready
A transaction is runnable only when every relation it touches is ready. This requires readiness-aware, disk-bounded scheduling and the out-of-order durable receipt mechanism from PR #9. Crashes during the COPY phase should continue resetting the whole base copy so the exported-snapshot correctness model is preserved.
Acceptance criteria
- Target apply progress advances before deferred secondary indexes finish.
- No normal long-running index build blocks active replay.
- Crash recovery safely reconciles interrupted concurrent index builds.
- Multi-table transactions remain atomic and ordered per relation.
- Partitioned, keyless, foreign-key, unique-index, and update-heavy workloads remain correct.
followbegins only after both catch-up and required schema finalization complete.- End-to-end tests prove source/target data and final index definitions match after cutover.
- Benchmarks show a material reduction in peak CDC backlog and catch-up duration.
- Lingua principale
- Go
- Stelle
- 6
- Fork
- 0
- Merge medio
- 28m
- PR unite (30g)
- 3
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di GetStream/pgmigrate
Tutte le issue di GetStream/pgmigrate
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
duplication
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
openvibely/openvibely#1443 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 80/100
keyxmakerx/Chronicle#1179 ·
I maintainer di solito rispondono entro 1 giorno
-
raised-by:worker
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
medici-finance/assay#2486 ·
I maintainer di solito rispondono entro 1 giorno
-
area/testing kind/bug triage/needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
cozystack/cozystack#4841 · 1 reazione ·
I maintainer di solito rispondono entro 2 giorni