Improve event delivery and ZeroMQ queue configuration
I maintainer di solito rispondono entro 1 giorno
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- java
- Ambito
- backend, distributed-systems
Direzione di ricerca
Start with RealtimeEventService.add() and work(), then trace how BlockEventLoad produces events and how queue pressure could pause and resume loading. Inspect NativeMessageQueue.start() in useNativeQueue mode, focusing on when SndHWM is configured relative to socket creation. Done means pending events apply backpressure without drops and sendQueueLength controls the publisher queue.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The event service has two issues affecting reliability and configurability:
RealtimeEventServiceuses an unboundedLinkedBlockingQueuewith an application-level 10,000-entry soft limit. Once that threshold is reached, new events are dropped instead of applying backpressure.NativeMessageQueuesets the send high-water mark (SndHWM) only after creating the socket, causing the configured send queue length to be ineffective.
We recommend introducing explicit backpressure between RealtimeEventService and BlockEventLoad, and correcting the timing of the SndHWM configuration.
Root Cause
1. RealtimeEventService drops events when the queue is full without applying backpressure
The queue is an unbounded LinkedBlockingQueue, and a single-threaded scheduler calls work() once per second to consume events in batches. The producer-side add() logic is:
public void add(Event event) {
if (queue.size() >= maxEventSize) { // maxEventSize = 10000
logger.warn("Add event failed, blockId {}.", ...);
return; // Drop the event directly
}
queue.offer(event);
}
When downstream consumption (work(), which runs once per second) cannot keep up with the event production rate and the queue reaches 10,000 entries, subsequent events are dropped directly, with only a warning being logged.
This is a drop-when-full strategy rather than a backpressure mechanism that applies pressure to the producer. Events generated while the queue is full are therefore silently lost.
2. NativeMessageQueue sets SndHWM too late, causing the configuration to be ineffective
In start(), the PUB socket is created first, and context.setSndHWM(sendQueueLength) is called afterward.
In ZeroMQ (JeroMQ), ZContext.setSndHWM only affects sockets created after the context-level setting is applied. The high-water mark is applied to the socket when it is created, so changing the context value afterward does not affect an already-created publisher socket.
As a result, the configured sendQueueLength is silently ignored, and the publisher continues to use the default high-water mark of 1000.
Reproduction
1. RealtimeEventService
Slow down downstream event consumption so that the event production rate remains higher than the consumption rate of once per second.
Once the queue reaches 10,000 entries, observe that subsequent events are dropped, with only an Add event failed warning being logged. The dropped events cannot be recovered.
2. NativeMessageQueue
In useNativeQueue mode, configure an explicit sendQueueLength. After startup, inspect the actual send high-water mark of the publisher.
The actual value remains at the default of 1000 instead of the configured value.
Impact
RealtimeEventService: When downstream consumption slows down, events beyond the 10,000-entry queue limit are silently dropped, causing real-time event loss and potentially resulting in subscribers missing events. The current drop-when-full + log strategy does not propagate backpressure to upstream producers.NativeMessageQueue: The configured send queue length is ineffective, preventing operators from controlling the publisher's send queue through configuration. When downstream consumption slows down and pending messages exceed the default high-water mark, excess messages may be dropped directly by the PUB socket, again resulting in silent event loss. The affected scope is limited touseNativeQueuemode.- Both issues affect the reliability and configurability of event delivery. They do not affect consensus or asset security.
Suggested Fix
-
RealtimeEventService: Introduce an explicit backpressure mechanism, such as anisBusy()method. When the number of pending events reaches the configured threshold,BlockEventLoadshould pause loading new events and resume once the queue size falls below the recovery threshold. This prevents events from being dropped when the queue reaches a fixed limit. -
NativeMessageQueue: Move thesetSndHWMconfiguration to before the socket is created, or directly callsetSndHWM(sendQueueLength)on the socket itself, so that the configured value takes effect.
- Lingua principale
- Java
- Stelle
- 4.2k
- Fork
- 1.8k
- Merge medio
- 3g 21h
- PR unite (30g)
- 13
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di tronprotocol/java-tron
-
type:feature
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
tronprotocol/java-tron#7013 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
[Feature] Validate ECKey inputs and improve key handlingForse già presa @Federico2014 l’ha presa 17 giorni fa. Apertatype:feature
Difficoltà 4/5 3-5 giorni Idoneità per principianti 40/100
tronprotocol/java-tron#6994 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
type:feature
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
tronprotocol/java-tron#6989 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Robustness fixes for multiple issuesForse già presa @xxo1shine l’ha presa 19 giorni fa. Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
tronprotocol/java-tron#6969 · 8 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
[Feature] decouple json-rpc filter processing from ManagerForse già presa @0xbigapple l’ha presa 18 giorni fa. Apertatype:feature
Difficoltà 5/5 Più di una settimana Idoneità per principianti 48/100
tronprotocol/java-tron#6963 · 8 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di tronprotocol/java-tron
Issue simili
-
[Bug] The producer summary counts an unreported client version as a second version and warns about a version mixForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
apache/rocketmq-dashboard#6110 ·
I maintainer di solito rispondono entro 4 giorni
-
`Processing lsp` never exits and leaves orphaned processesForse già presa @overcast302 l’ha presa oggi. Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
processing/processing4#1578 · 1 commento ·
-
ASM is not up-to-dateAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 60/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] S3 CORS responses omit Access-Control-Allow-Credentials for matched originsForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
floci-io/floci#5369 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
securityHeaders replaces a route's own Content-Security-Policy (0.9.9; weakens embedders' pages)Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno