A QUEUED batch job with no jobs cannot be finished or deleted, and holds a queue slot
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
Research direction
Look at the batch job and job models in the compute job manager backend, likely in a directory like app/models/. Find the validation logic for finishing a batch job and the queue slot management. The issue is about state transitions when a QUEUED batch has zero jobs; you need to understand the lifecycle and constraints. A fix could be in the batch job finish endpoint or an automatic cleanup trigger. Test by reproducing the bug with the provided Python script.
Written by the indexing model from the issue text.
Description
Bug report
Describe the bug
A batch job can be driven into a state that no client can leave, and that occupies one of the
batchjobs_per_queue_limit slots indefinitely:
- A batch job sits in
QUEUEDand is never scheduled (see the separate note below on why). - The user deletes its jobs with
DELETE /jobs/{id}- all of them were inPLANNEDstate, so
no results are lost - expecting the queue slot to be released. - The batch job is now left with
job_ids: []butstatus: QUEUED.
From that state:
| Attempt | Result |
|---|---|
PATCH /batch_jobs/{id}/finish |
422 - ValidationError(loc=[], msg="", type="IntegrityError") |
POST /jobs with batch_job_id set to the empty batch |
422 - Batch job has invalid status: queued |
DELETE /batch_jobs/{id} |
does not exist in the OpenAPI spec |
PATCH /batch_jobs/pop, /peek |
403 (device-only, correctly) |
So the platform cannot finish it either (its own integrity constraint rejects a zero-job batch),
and the client has no way to remove it. With batchjobs_per_queue_limit = 5, five such batches
lock the account out of the backend completely. The only symptom the user sees is:
HTTP 429 detail="Backend type Tuna-17 allows a maximum of 5 batch jobs in the queue at the
same time"
which does not name the offending batches or even hint that they are empty.
To reproduce
import asyncio, compute_api_client as c
from qi2_shared.client import config
async def main():
async with c.ApiClient(config()) as cl:
bj = await c.BatchJobsApi(cl).create_batch_job_batch_jobs_post(
c.BatchJobIn(backend_type_id=7))
await c.BatchJobsApi(cl).enqueue_batch_job_batch_jobs_id_enqueue_patch(bj.id)
# ... create one job in it, then:
# await c.JobsApi(cl).delete_job_jobs_id_delete(job_id)
# now bj is QUEUED with job_ids == []
asyncio.run(main())
Observed batches
On 2026-09-22 five batches reached this state and held all five queue slots for about two hours:
844305 queued 07:18:02 UTC
844307 queued 07:31:01 UTC
844309 queued 07:57:30 UTC
844310 queued 08:09:49 UTC
844311 queued 08:21:28 UTC
They were cleared some time later without any action on our side, so this is not urgent for us -
but the state itself is reachable by any user with two documented calls.
Expected behaviour
Either of these would be enough:
- a
DELETE /batch_jobs/{id}endpoint, or - automatic cleanup of a
QUEUEDbatch whose last job is removed, or finishaccepting a zero-job batch so the state is at least terminable from the client.
Nice to have: have the 429 name the batches occupying the slots, and/or expose a queue position.
Why the batch was never scheduled (related, and our mistake, not a platform bug)
Worth recording because it is easy to hit and hard to diagnose. The wait argument counts cycles
of 20 ns, and per the operational-specifics documentation the idle time is paid once per shot, so a
circuit's execution time scales with shots x delay. job_execution_time_limit (300 s for Tuna-17)
applies to the whole batch. A single circuit using wait(65536) needs roughly 430 s at 65536 shots,
so a batch containing it can never fit the budget - it just never runs. One circuit per batch, and
keeping the delay within the budget, resolves it.
- Dominant language
- Python
- Stars
- 0
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from QuTech-Delft/compute-api-client
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
All issues in QuTech-Delft/compute-api-client
Similar issues
-
enhancement good first issue Stellar Wave trivial
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
StellarCanary/ProtocolCanary-Fixtures#258 ·
Maintainers usually reply within 1 day
-
github_actions
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Hochfrequenz/aibap.mcp#578 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
mishraprafful/multihull#150 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
python-caldav/caldav#735 ·
Maintainers usually reply within 1 day