Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Set per-table autovacuum and analyze thresholds on large tables

Aperta
#6,079 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
35/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
django, postgresql, python

Direzione di ricerca

Inizia esaminando le convenzioni delle migration e la procedura di deploy in Makefile:32-40, quindi individua l’entry point del management command utilizzato per le operazioni sul database. Implementa le impostazioni reversibili della tabella, il comportamento di errore in caso di lock timeout e la validazione dell’argomento della tabella descritti qui. Il lavoro è completato quando la migration, il comando ANALYZE, l’invocazione di deploy e i test elencati soddisfano i criteri di accettazione.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

DEV: backend

❌ This issue is not open for contribution. Visit Contributing guidelines to learn about the contributing process and how to find suitable issues.

Overview

  • Default autovacuum_vacuum_scale_factor = 0.2 defers vacuum until 20% of a table's rows are dead — 3.36M tuples on contentnode, 11.9M on assessmentitem.
  • Neither table has ever been vacuumed or analyzed, leaving the planner without statistics for the largest table in the database.

Complexity: Low
Target branch: hotfixes

Context
  • ALTER TABLE ... SET (autovacuum_*) is metadata-only and does not rewrite the table, but takes a brief ACCESS EXCLUSIVE lock. On contentnode that can queue behind a long-running transaction and block everything behind it — gunicorn's timeout is 4000s, so long transactions are possible.
  • pg_class.reloptions is currently null for all three tables, and Django exposes no model-level API for storage parameters.
  • The app image has no psql, so ANALYZE must be issued through a database cursor.
The Change
  • A migration should set per-table autovacuum and analyze thresholds on contentnode, assessmentitem, and file: autovacuum_vacuum_scale_factor = 0.01 and autovacuum_analyze_scale_factor = 0.005.
  • file should additionally get autovacuum_vacuum_insert_scale_factor = 0.05 — at 113M insert-heavy rows the dead-tuple threshold alone never fires, but the visibility map still needs maintaining.
  • The migration should set a lock_timeout and fail rather than retry, so a blocked ACCESS EXCLUSIVE acquisition does not queue readers behind repeated attempts.
  • A general-purpose management command should run ANALYZE against tables named at invocation, issued through a database cursor.
  • make deploy-migrate should invoke that command for the three tables, per the procedure at Makefile:32-40.
Acceptance Criteria
General
  • pg_class.reloptions on contentcuration_contentnode, contentcuration_assessmentitem, and contentcuration_file contains autovacuum_vacuum_scale_factor=0.01 and autovacuum_analyze_scale_factor=0.005
  • contentcuration_file additionally has autovacuum_vacuum_insert_scale_factor=0.05
  • The migration reverses cleanly, resetting all three tables to no storage parameters
  • The migration aborts with a non-zero exit when it cannot acquire its lock within lock_timeout
  • A management command runs ANALYZE against tables passed as arguments
  • make deploy-migrate invokes that command for the three tables
Testing
  • last_analyze is non-null on all three tables after the deploy step runs
  • last_autovacuum becomes non-null on contentnode within 24 hours of the thresholds applying
  • Unit test covers the command's table-argument handling and its error on an unknown table

AI usage

Drafted with Claude Code, which measured the vacuum state, analyze state, and row counts cited here through read-only queries against the production database. I confirmed the deploy-migrate procedure and migration precedent against the Studio source, and set the threshold values and the fail-loudly behaviour.

Lingua principale
Python
Stelle
191
Fork
308
Merge medio
2g 6h
PR unite (30g)
65

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di learningequality/studio

Tutte le issue di learningequality/studio

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.