Workflow dep resolution leaves `scheduled_at` stale, breaking queue delay monitoring
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- go, postgresql
- Área
- backend, databases, observability
Línea de trabajo
Locate the WorkflowStageJobs and WorkflowStageJobsByIDMany queries and inspect the jobs_to_make_available CTE plus its UPDATE of river_job. Confirm the behavior for jobs becoming available versus remaining scheduled, then choose and implement the agreed timestamp approach; done means dependency-resolved available jobs report an accurate queue-delay timestamp without breaking scheduled jobs.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Description
When WorkflowStageJobs / WorkflowStageJobsByIDMany resolve dependencies and transition a job from pending to available, the scheduled_at column is not updated. It retains its original value from insertion time, which can be hours or months old for long-running workflows.
The UPDATE in both queries only sets state and metadata.workflow_staged_at:
UPDATE river_job
SET
state = jobs_to_make_available.new_state,
metadata = jsonb_set(metadata, '{workflow_staged_at}'::text[], $1::jsonb, true)
FROM jobs_to_make_available
WHERE river_job.id = jobs_to_make_available.id
The jobs_to_make_available CTE already reads scheduled_at to decide the target state (available if scheduled_at <= now() + 5s, otherwise scheduled), so by the time the UPDATE executes, the original scheduled_at value has served its purpose.
Impact
Any monitoring that uses NOW() - scheduled_at on available jobs to measure queue delay will report wildly inflated values for dependency-resolved workflow jobs. For workflows where deps take hours or months to resolve, this produces false alarms on queue health metrics.
Current workaround
We discovered that workflow_staged_at is already stamped in metadata during dep resolution, so we use it as a fallback in our metrics query:
MAX(
CASE
WHEN metadata ? 'workflow_staged_at'
THEN NOW() - (metadata->>'workflow_staged_at')::timestamptz
ELSE NOW() - scheduled_at
END
) as oldest_delay
This works but requires casting a JSONB string to timestamptz in an aggregate query, which is less ergonomic than using the native scheduled_at column directly.
Proposed solutions
Either of these would address the problem:
-
Update
scheduled_at = now()inWorkflowStageJobs/WorkflowStageJobsByIDManywhen transitioning jobs toavailable. This makesscheduled_ataccurately reflect when the job became eligible for pickup, consistent with how non-workflow jobs behave. For jobs transitioning toscheduled(because theirscheduled_atis still in the future), no change is needed —scheduled_atis already correct. -
Add a first-class
available_atcolumn toriver_jobthat records when a job entered theavailablestate, regardless of how it got there (direct insert, scheduled time reached, or workflow dep resolution). This would give monitoring queries a reliable, indexed timestamp without relying onscheduled_atsemantics or JSONB metadata. It would also benefit non-workflow use cases like jobs inserted withPending: truethat are later moved toavailableby application code.
Environment
- River Pro v0.22.0
- PostgreSQL
- Lenguaje dominante
- Go
- Estrellas
- 5.7k
- Forks
- 187
- Merge medio
- 2 d 17 min
- PR fusionados (30 d)
- 43
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de riverqueue/river
-
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
riverqueue/river#1411 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 45/100
riverqueue/river#1358 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
River job stuck at runningAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
riverqueue/river#1258 · 7 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
riverqueue/river#1225 · 14 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
riverqueue/river#1183 · 2 reacciones ·
Los mantenedores suelen responder en 1 día
Todos los issues de riverqueue/river
Issues similares
-
area/tests theme/ci-dx
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
bug needs triage pkg/translator/faro
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
open-telemetry/opentelemetry-collector-contrib#51759 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
appbaseio/reactivesearch-api#402 ·
-
GET /api/v1/system/api_keys returns the full API token in plaintextPosiblemente ocupada @Harsh23Kashyap la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
infiniflow/ragflow#20555 · 1 reacción ·
Los mantenedores suelen responder en 1 día