Add mechanism to increase efficiency of reconnection process after upgrade
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- rust
Línea de trabajo
Comienza volviendo a ejecutar la prueba de reconexión mencionada con permanent_error_backoff establecido en 1 segundo y, después, rastrea el comportamiento de la lista de no llamar, known_addresses y la clasificación de conexiones descrito en el escenario. El issue enumera varios enfoques posibles, pero no selecciona ninguno ni identifica archivos; por lo tanto, los criterios de finalización requieren que los maintainers acuerden el diseño y una mejora fiable de la reconexión después de reiniciar los nodos.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
In 1.6 and onward there is a case resulting in a higher-than-optimal time or number of attempts to reestablish a connection. This was found in testing:
The minor potential issue is timing in general - tests use fast block times and node-kill-timers, if I am not mistaken? If so, the backoff timers might be too long to allow the system to work reliably. What you may be running into is what I call the “upgrade problem”, since I expect the same scenario to potentially occur on an upgrade: When a node is killed/restarted, its node ID changes randomly as well as connections timing out. Here’s the scenario from Alice’s perspective where we go from Alice > Cecil to Cecil > Alice:
- Alice loses connection to Cecil.
- Alice fails to reconnect immediately while Cecil boots.
- Eventually, within the reconnect delay, Cecil boots up.
- Alice is in Cecil’s known_addresses, so Cecil connects to Alice.
- An incoming connection is set up, all is good.
- Alice connects to Cecil, notices the low connection rank, blocks Cecil’s socket address (for outgoing) for 10 minutes.
This is the “good” case, since the connection rank inverted, we were able to establish a connection quickly. However, what if Cecil goes down again after 1 minute and we flip the connection rank again, so that Alice > Cecil again?
- Alice loses the incoming connection from Cecil. It does not care, since it is an inbound connection.
- Cecil reboots, establishes a connection to Alice due to known_addresses.
- Both notice the low rank of the connection, but Alice learns of Cecil’s address.
- However, this time around, when learning Cecil’s address it is still on the do-not-call list.
- For another 9 minutes, we cannot establish a connection between these two nodes.
You can verify this by running the test again with permanent_error_backoff set to 1 or 5 seconds, I recommend doing that. Potential fixes:
- Remove entries from the do-not-call list if the addresses are given to us by Magic Mike. This is a potential security issue though, as we can never verify the connection between a NodeId and a SocketAddr. A clever attacker could cause a node to get itself firewalled from every good validator this way.
- Lower the permanent_error_backoff -- not great, since it increases reconnection spam.
- Persist the NodeId of a node across restarts. This is likely the best approach and what we had 3 years ago, but more work.
- Implement the gossip-less address syncing that I have in mind, to avoid having to have such long backoff timers.
- Ignore it, since it will only present itself in arbitrary testing situations (unless a node is in a crash loop).
The achilles heel of the system is the fact that we need to connect to unverified socket addresses. The original system used to sign SocketAddr, NodeId pairs for that reason.However, I encourage you to test with permanent_error_backoff set to 1 second to test this hypothesis at least.
- Lenguaje dominante
- Rust
- Estrellas
- 397
- Forks
- 224
- Merge medio
- 13 d 23 h
- PR fusionados (30 d)
- 1
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de casper-network/casper-node
-
consensus networking
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
casper-network/casper-node#1922 ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
casper-network/casper-node#5449 ·
-
Proposal for a New Build Numbering SchemeQuizá libre de nuevo @sacherjj la tomó hace 425 días y no hay ningún pull request abierto. Abierto
casper-network/casper-node#5296 · 2 comentarios · 1 asignado ·
-
casper-network/casper-node#5273 · 2 comentarios · 1 asignado ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 45/100
casper-network/casper-node#5259 · 5 comentarios ·
Todos los issues de casper-network/casper-node
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
stellar/stellar-cli#2773 ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
voidzero-dev/oxc-angular-compiler#511 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 86/100
yantrikos/yantrik-os#539 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
documentation station:mac ui-dashboard
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
rolter-ai/rolter#2490 · 1 comentario ·
Los mantenedores suelen responder en 1 día