[Follow up] Time management after #8416
Maintainers usually reply within 1 day
@aldbr is already working on this.
Since Aug 25, 2026.
Assessment
This issue has not been assessed yet.
Description
Bug Description
/LocalSite/CPUTimeLeft is the CPU work left in the batch slot. The JobAgent publishes it every cycle, and three things read it: the Matcher, through the CPU work the CE advertises; the Watchdog, which counts it down; and the payload itself, when it is elastic and has to decide how much work to take on.
But the payload never gets all of it. Once the payload stops, the JobWrapper still has to upload the outputs and the logs, and the Watchdog reserves StopMargin (300 s by default) at the far end for exactly that — it stops the payload while that much is still on the clock. Nothing upstream of the Watchdog knows. So a job whose declared CPUTime just fits the slot is matched into it, and an elastic payload that sizes itself against /LocalSite/CPUTimeLeft commits to work it will not be allowed to finish. Both are then
stopped part-way.
Being stopped is itself lossy. The Watchdog's only response when the budget runs out is self.spObject.killChild(), i.e. SIGTERM escalating to SIGKILL. A payload killed mid-unit loses everything it had produced but not yet written, and because the job then fails, none of it is uploaded or recorded either.
There used to be a way out of that. The Watchdog still parses StopSigRegex, StopSigNumber, StopSigStartSeconds and StopSigFinishSeconds from the JDL, and used to signal the payload to wind down before killing it. That code is gone.
Finally, an application that does stop on a signal exits 128 + N — 130 for SIGINT, 138
for SIGUSR1. JobWrapper.postProcess treats any non-zero exit as an application error, so
a payload that did exactly what it was asked would be recorded as a failure.
Steps to Reproduce
- Configure a queue whose slots are comparable to the length of one job (or let a pilot fill
until little is left). - Submit a job whose JDL
CPUTimeis close to/LocalSite/CPUTimeLeft, or an elastic
payload that reads/LocalSite/CPUTimeLeftand sizes its work from it. - Watch it be matched, run, and be killed by the Watchdog with
StopMarginstill to go.
The elastic case is the one that shows the sizing error clearly, because the payload commits
to a specific amount of work up front. LHCb MC is the example we hit it with, but nothing
about the defect is VO-specific: any payload that sizes itself against the advertised slot
over-commits by StopMargin.
Expected Behavior
- The CPU work advertised to the Matcher, and published in
/LocalSite/CPUTimeLeft, is what
a payload may actually consume — the post-processing reserve already deducted, once, by
whoever publishes it. - A payload that knows how to wind down can be told to, early enough to finish its current
unit of work and write its output, rather than only ever being killed. - A payload that stops when asked is not recorded as an application error.
Actual Behavior
Observed on an elastic LHCb MC job (job 1473012931), matched on
cycle 6 of 10 with 1357 s of wall clock left in the slot:
CPUTimeLeft = 37878 (normalized units) CPUNormalizationFactor = 27.9
CPUTime = int(37878 / 27.9) = 1357 s
eventsToProduce = int(floor(1357 * 27.9) / 154) = 245
willProduce = int(245 * 0.75) = 183 # VO safety factor
The payload was sized for 183 units against 1357 s, but the Watchdog only ever intended to
let it have 1057. It produced 150, was killed, and the 4.9 MB output file it had written was
never uploaded. The job is Failed, so it also never reaches the Bookkeeping — which means
failures of this kind cannot feed back into the per-unit cost estimate that sized them.
Environment
- DIRAC integration
- Payload: elastic, sizes its own work from /LocalSite/CPUTimeLeft
Relevant Log Output
Job has reached the CPU limit of the queue wallClockLeft=297s
'FinalMinorStatus': 'Job has reached the CPU limit of the queue',
'ExecTime': 1065, 'ProcessedEvents': 0
ProcessedEvents: 0 despite 150 having been produced: the payload was killed before it could
report them.
Additional Context
Proposed fix, in two parts. They are independent by construction and can be reviewed
separately:
- #8528: report a watchdog-stopped payload for what it is, rather than as
"No outputs generated from job execution", and treat128 + Nas a clean exit whenN
is the signal the Watchdog itself sent. The second half is a no-op until the graceful stop
exists (stopSigSentis never set today), so this can go first on its own. - next PR: deduct
StopMarginonce, inJobAgent.initialize(), so the Matcher and
the payload both see a budget they can actually spend; and restore the graceful stop on
that same budget, withStopSigRegexmatching the command line again as it did before
0e67f781de.
Related:
- #8416 the single-writer design this relies on: the JobAgent publishes
/LocalSite/CPUTimeLeftand everyone else reads it, because a containerised payload cannot
reach the batch system to recompute it. - #8346 draining pilots; its second point, jobs that produce nothing when stopped
and are hard to tell from real failures, is the same complaint from the other end.
- Dominant language
- Python
- Stars
- 126
- Forks
- 191
- Avg merge
- 1d 8h
- Merged PRs (30d)
- 32
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from DIRACGrid/DIRAC
-
[Bug]: Docs: always false condition in tornado run file?Possibly taken @andresailer claimed this 4 days ago. OpenBug
DIRACGrid/DIRAC#8807 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
Drop boto in all DIRAC reposMay be free again @natthan-pigoux claimed this 65 days ago, and no pull request is open. Open
DIRACGrid/DIRAC#8641 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
Create a legacy adaptor for PilotManagerPossibly taken @AcquaDiGiorgio claimed this 24 days ago. Open
DIRACGrid/DIRAC#8597 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
DIRACGrid/DIRAC#8549 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
mikf/gallery-dl#9791 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
fossasia/eventyay#6151 · 1 comment ·
Maintainers usually reply within 1 day
-
P4: low tooling
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
jeffknupp/association#318 ·
-
azure-cost bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
microsoft/GitHub-Copilot-for-Azure#3330 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
raullenchai/Rapid-MLX#4097 ·
Maintainers usually reply within 1 day