rivetkit-core: a failed scheduled dispatch pass drops due one-shot actions and strands recurring jobs
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
- #5834 di @aryanghai12 — chiusa senza merge
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 22/100
Direzione di ricerca
Start with take_due_schedule_dispatches in rivetkit-rust/packages/rivetkit-core/src/actor/schedule.rs and the invariant at lines 736-737; read the contract in docs/actors/content/docs/schedule.mdx. Run the three regression tests in rivetkit-rust/packages/rivetkit-core/tests/schedule.rs to see the claimed one-shots vanish and the schedule_running marker leak. Done means the pass commits atomically and releases only its own markers on error. Note a PR is already open and the failure semantics are still under discussion.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
ActorContext::take_due_schedule_dispatches claims every due one-shot schedule by deleting it in one committed batch. It then advances each due recurring schedule in its own committed batch.
If anything fails after the one-shot claim commits, the function returns Err and the collected dispatches are dropped.
This can cause:
- due one-shot actions to be deleted and never attempted
- recurring jobs advanced earlier in the same pass to lose that run
- their history row to remain
running - their
schedule_runningmarker to remain set, causing later runs to be recorded asskipped
This is separate from the crash window inherent to claim-then-dispatch. The issue here occurs while the process remains alive and the function returns an ordinary error.
The behavior is also inconsistent with the scheduling contract in docs/actors/content/docs/schedule.mdx:
- schedules are documented as surviving sleep, restarts, upgrades, and crashes
- a failed one-shot is only considered complete after its attempted invocation
- overlapping recurring runs are the cases that should be recorded as
skipped
There is also an in-code invariant at rivetkit-rust/packages/rivetkit-core/src/actor/schedule.rs:736-737 stating that already-claimed work must not be discarded by a transient failure.
Reproduction
The failure can be reproduced with deterministic regression tests in:
rivetkit-rust/packages/rivetkit-core/tests/schedule.rs
The tests cover:
failed_dispatch_pass_does_not_lose_due_one_shotsfailed_dispatch_pass_does_not_strand_recurring_jobsunparsable_recurring_row_does_not_lose_one_shots_or_strand_other_jobs
The first two inject a SQLite failure during the recurring-history write, after the one-shot claim has already committed.
The third exercises a stored recurring row whose cron expression can no longer be parsed after the due rows have been loaded and the one-shot claim has committed.
On the current main commit, the observed behavior is:
- the due one-shot is deleted and never dispatched
- a recurring job advanced earlier in the pass keeps a
runninghistory entry without a dispatch - the leaked
schedule_runningmarker causes later runs of that job to be recorded asskipped
Root cause
The dispatch pass is not atomic.
The current flow is effectively:
- claim due one-shots and commit
- process recurring rows one by one
- advance and record each recurring row in a separate transaction
- return
Erron a later failure
Once step 1 commits, the function can no longer roll that work back.
Proposed direction
One possible fix is to:
- plan all claims, recurring advances, and history writes first
- commit them in a single
execute_batch - build and publish dispatches only after a successful commit
- release only the
schedule_runningmarkers created by the current pass on any error path
This would keep the current error-return behavior while making the dispatch pass all-or-nothing.
No SQL schema or protocol changes should be required.
I'm happy to open a PR for the fix once the preferred failure semantics are confirmed.
Questions
- Is making the whole dispatch pass atomic the preferred approach, or would you rather isolate failures per recurring job?
- Should a stored schedule that can no longer be parsed be quarantined rather than causing every dispatch pass to fail?
- Lingua principale
- Rust
- Stelle
- 6.3k
- Fork
- 271
- Merge medio
- 2g 16h
- PR unite (30g)
- 89
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Nessuna guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di rivet-dev/rivet
-
Rust client drops connection-level errors when actionId is null (JSON/CBOR)Forse già presa @DibbayajyotiRoy l’ha presa 6 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
rivet-dev/rivet#5819 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Queue docs say completable messages are removed on receive, but they stay stored until complete()Forse già presa @Saisharathchandranandnetha l’ha presa 17 giorni fa. Aperta
Difficoltà 1/5 1-3 ore Idoneità per principianti 91/100
rivet-dev/rivet#5762 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Serverless listener forwards /start to the application when basePath is "/"Forse già presa @Saisharathchandranandnetha l’ha presa 20 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
rivet-dev/rivet#5755 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
rivet-dev/rivet#5850 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
rivet-dev/rivet#5837 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di rivet-dev/rivet
Issue simili
-
defect
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 2 giorni
-
enhancement user-priority/P3
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
I maintainer di solito rispondono entro 1 giorno