Monitoring crashed workers and rescuing stuck jobs

Abierto
#773 1 comentario 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
25/100
Tipo de issue
Nueva funcionalidad
Claridad
Necesita aclaración
Estado de actividad
Estancado
Stack tecnológico
go

Línea de trabajo

Start by reviewing the rescuer, worker, leader-election, and global rescue-timeout behavior described in the issue. Compare the proposed worker healthcheck timestamps with the current rescue model and identify failure cases and design constraints. Done means the project has an agreed approach and clear behavior for detecting failed workers and rescuing their jobs.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Hello there!

After reading about leader election, updating timestamps, my initial thought was that rescuer would also monitor all workers that have running jobs attached to them.

I assumed that all workers would update some kind of a healthcheck timestamp every 5 seconds or so. And if some worker hasn't updated their timestamp in a reasonable time period, all its currently running jobs should be rescued.

However, I see that jobs are only rescued based on a global rescue timeout.

What do you think about this?

Lenguaje dominante
Go
Estrellas
5.7k
Forks
179
Merge medio
15 h 43 min
PR fusionados (30 d)
13

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de riverqueue/river

Todos los issues de riverqueue/river

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.