Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Sidecar metrics collection over plugin gRPC has no deadline — one stuck operation hangs the whole instance /metrics

Aperta
#1,045 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
55/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
go, grpc

Direzione di ricerca

Traccia il percorso di collect delle metriche gRPC del plugin dell’instance sidecar e riproduci il problema con un’operazione del plugin bloccata, quindi esamina come l’endpoint /metrics gestisce gli errori del collector. Verifica che un timeout per collect consenta all’endpoint di restituire le metriche rimanenti e registrare un errore quando la chiamata al plugin si blocca, mentre le metriche del plugin funzionanti continuano a funzionare.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug

Environment

CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0, PostgreSQL 17/18 clusters, VersityGW (posix) S3 endpoint.

What happens

The instance sidecar collects plugin metrics over the plugin gRPC with no deadline on the collect path. Consequence: if any plugin operation gets stuck (in my case a retention/catalog-maintenance run that never completes against a posix-backend S3 gateway — filed separately), the next scrape blocks on the collector and the entire /metrics endpoint of that instance hangs indefinitely. Prometheus marks the target down and every metric of that instance disappears, not just the plugin's own collectors.

Observed in production: the primary instance's exporter dark for 6.7 hours while the database itself was healthy — WAL-archiver and backup-staleness alerting was blind precisely on the instance where it matters.

Expected

  • A per-collect deadline (a few seconds) on the plugin metrics gRPC call.
  • On timeout/error: return the rest of the metrics plus an error counter (e.g. ..._collector_errors_total), instead of hanging the whole endpoint. A misbehaving plugin should degrade its own metrics, never the instance's.

Reproduction sketch

  1. ObjectStore against VersityGW (posix backend) with any retentionPolicy set.
  2. Wait for a retention run to start (~30 min cadence) — it never completes against this backend.
  3. Scrape the instance sidecar's metrics port: the request hangs until the client gives up; up goes to 0 for the instance.

Companion issue describes the retention hang itself; this one is about the missing deadline that turns any such hang into a full metrics outage.

Lingua principale
Go
Stelle
198
Fork
81
Merge medio
2g 6h
PR unite (30g)
19

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di cloudnative-pg/plugin-barman-cloud

Tutte le issue di cloudnative-pg/plugin-barman-cloud

Issue simili

Altre issue su Go

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.