Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Driver-level rx-drop counters (kernel-drops, ppline-drops) aren't exposed via the Prometheus endpoint

Aperta
#1,790 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
35/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Attiva
Stack tecnologico
prometheus, rust
Ambito
cli, observability

Direzione di ricerca

Inizia dall’output dello stato del driver di dataplane-cli e dall’endpoint /metrics di Prometheus, quindi confronta le metriche vpc_* esistenti con kernel-drops e ppline-drops. Esamina i comandi correlati, inclusi show vpc peerings/routing/summary e show pipeline stats o stages, insieme all’indagine in githedgehog/fabricator#1989. Il lavoro è completato quando l’ambito delle metriche e la semantica di kernel-drops sono stati decisi e l’esposizione richiesta è chiaramente specificata.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Found while investigating githedgehog/fabricator#1989. Two drop signals live on opposite sides of a gap, each visible through exactly one interface:

  • dataplane-cli show driver status exposes per-AF_PACKET-worker kernel-drops and ppline-drops. Nothing on /metrics corresponds to it. Full metric list from a live gateway is 8 names, all vpc_*:
    vpc_byte_count, vpc_byte_rate, vpc_packet_count, vpc_packet_rate,
    vpc_pair_drops_byte_count, vpc_pair_drops_byte_rate,
    vpc_pair_drops_packet_count, vpc_pair_drops_packet_rate
    
  • Conversely, vpc_pair_drops_packet_count (metrics-only) has no CLI equivalent: show vpc peerings/routing/summary don't surface a per-pair drop count, and show pipeline stats/show pipeline stages (the two commands that look like the natural home for it) both currently answer Not supported: Not implemented yet.

This matters because production monitoring is the Prometheus scrape, not an interactive dataplane-cli session. In the #1989 investigation, kernel-drops turned out to be the one signal that correlated with the failure (bursty, clustered on the exact NAT tests driving real throughput, on the one or two workers actually carrying that traffic, silent otherwise) - and it would have been invisible to any dashboard or alert built on the existing metrics, discoverable only by someone thinking to run the CLI by hand.

Ask: is there a reason kernel-drops/ppline-drops aren't exposed as metrics today (cardinality, overhead, or just not gotten to), or would it be reasonable to add them (per worker, or aggregated) so a burst-driven drop shows up alongside vpc_pair_drops_packet_count in whatever already watches that?

Separately, confirming the semantics would help interpret this and future readings correctly: does kernel-drops mean what its name suggests, packets the kernel dropped before this worker's AF_PACKET socket could read them (e.g. ring full)?

Lingua principale
Rust
Stelle
20
Fork
9
Merge medio
4g 9h
PR unite (30g)
35

Preparare l'ambiente

  • Include un Dockerfile o un file Docker Compose
  • Nessun modello di pull request
  • Nessuna guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di githedgehog/dataplane

Tutte le issue di githedgehog/dataplane

Issue simili

Altre issue su Rust

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.