SimulatePubnetMixedLoad: a node that loses sync can never rejoin (local-only archive), failing the mission
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- fsharp
- Ambito
- distributed-systems, testing-qa
Direzione di ricerca
Inizia da StellarNetworkData.fs intorno alla riga 580 e dalla configurazione di FullPubnetCoreSets usata da SimulatePubnetMixedLoad. Riproduci la missione mentre esamini i log di catchup e la configurazione dell’archivio locale. Il lavoro è completato quando un nodo che perde la sincronizzazione può eseguire il catchup e tornare a Synced!, consentendo il superamento del controllo finale fullCoreSet di EnsureAllNodesInSync.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Observed on a SimulatePubnetMixedLoad run (stellar-core 27.1.1 133f5bcf0, protocol 27). Pod ssc-0942z-2aa52f-sts-node-139-0 ended the run stuck in catchup, so the mission's final EnsureAllNodesInSync fullCoreSet failed.
What the node's log shows, in order:
- It's a non-tier1 leaf with a single peer for the whole run —
TARGET_PEER_CONNECTIONS: 1,PREFERRED_PEERS_ONLY. It spends the first ~2.5 min unable to resolve its preferred peer and rejecting all others, then authenticates one connection:
Unable to resolve peer 'ssc-...-sts-satoshipay-2...': Host not found (authoritative)
Non preferred inbound authenticated peer 172.22.177.113:11625 rejected because all available slots are taken.
Authenticated to 172.22.177.113:11625
- Under load (~1150 tx/ledger, network delay enabled) tx-set fetches from that one peer degrade until it drops out of sync. Apply time stays at 100–400 ms/ledger throughout, so it's starved for data, not slow to apply:
'fetch-e65310' took 13.515000 s
Ledger took 25.794006148 seconds (ledger 116)
'fetch-3cd058' took 40.518002 s
Herder WARNING Lost track of consensus
Lost sync, local LCL is 123, network closed ledger 126
- Catchup then can't proceed.
FullPubnetCoreSetssetshistoryNodes = Some([])(StellarNetworkData.fs:580), so the only archive is the node's ownlocalat/data/history. It last published checkpoint 63 before losing sync, so the checkpoint it now needs was never written, and it retries indefinitely:
Catching up to ledger 127: Downloading state file history/00/00/00/history-0000007f.json
cp: cannot stat '/data/history/history/00/00/00/history-0000007f.json': No such file or directory
Missing HAS for ledger 127: maybe stale archive local
That message repeats with growing backoff for the remaining ~11 minutes of the run. Meanwhile it keeps buffering externalized ledgers (mSyncingLedgers reaches 66) without ever applying them, and never returns to Synced!.
So in these missions a single loss of sync is terminal for that node, and the mission fails at the end regardless of how the rest of the network behaved.
- Lingua principale
- F#
- Stelle
- 10
- Fork
- 22
- Merge medio
- 1g 23h
- PR unite (30g)
- 4
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di stellar/supercluster
-
bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 48/100
stellar/supercluster#432 ·
-
Failed catchup ranges are retried only after all other ranges finish, adding hours to runtimeForse già presa @sisuresh l’ha presa 69 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
stellar/supercluster#409 ·
-
MinBlockTimeTest hardeningAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
stellar/supercluster#401 · 1 commento ·
-
Post-mission DumpData/GetRawMetrics crashes when cluster does not resolve .local TLDForse già presa @tomerweller l’ha presa 112 giorni fa. Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 66/100
stellar/supercluster#399 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
stellar/supercluster#397 ·
Tutte le issue di stellar/supercluster
Issue simili
-
[Bug] sglang.Engine `schedule_policy="priority"` is an advertised choice that always kills the schedulerForse già presa @charan-rathore l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
sgl-project/sglang#43161 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Pluggable component responses larger than 4 MiB fail with ResourceExhaustedForse già presa Una pull request collegata a questa issue è aperta o già unita. Apertakind/bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
ScalingModifiers formula fails with "formula returned non-float result" when expression evaluates to an integerForse già presa @Sarthak-Pandey l’ha presa oggi. Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
kedacore/keda#8270 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
area:manager bug triage:confirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Cotal-AI/Cotal#3404 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno