Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

A QUEUED batch job with no jobs cannot be finished or deleted, and holds a queue slot

Open
#5 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
45/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
api, backend

Research direction

Look at the batch job and job models in the compute job manager backend, likely in a directory like app/models/. Find the validation logic for finishing a batch job and the queue slot management. The issue is about state transitions when a QUEUED batch has zero jobs; you need to understand the lifecycle and constraints. A fix could be in the batch job finish endpoint or an automatic cleanup trigger. Test by reproducing the bug with the provided Python script.

Written by the indexing model from the issue text.

Description

Bug report

Describe the bug

A batch job can be driven into a state that no client can leave, and that occupies one of the
batchjobs_per_queue_limit slots indefinitely:

  1. A batch job sits in QUEUED and is never scheduled (see the separate note below on why).
  2. The user deletes its jobs with DELETE /jobs/{id} - all of them were in PLANNED state, so
    no results are lost - expecting the queue slot to be released.
  3. The batch job is now left with job_ids: [] but status: QUEUED.

From that state:

Attempt Result
PATCH /batch_jobs/{id}/finish 422 - ValidationError(loc=[], msg="", type="IntegrityError")
POST /jobs with batch_job_id set to the empty batch 422 - Batch job has invalid status: queued
DELETE /batch_jobs/{id} does not exist in the OpenAPI spec
PATCH /batch_jobs/pop, /peek 403 (device-only, correctly)

So the platform cannot finish it either (its own integrity constraint rejects a zero-job batch),
and the client has no way to remove it. With batchjobs_per_queue_limit = 5, five such batches
lock the account out of the backend completely. The only symptom the user sees is:

HTTP 429  detail="Backend type Tuna-17 allows a maximum of 5 batch jobs in the queue at the
                    same time"

which does not name the offending batches or even hint that they are empty.

To reproduce

import asyncio, compute_api_client as c
from qi2_shared.client import config

async def main():
    async with c.ApiClient(config()) as cl:
        bj = await c.BatchJobsApi(cl).create_batch_job_batch_jobs_post(
            c.BatchJobIn(backend_type_id=7))
        await c.BatchJobsApi(cl).enqueue_batch_job_batch_jobs_id_enqueue_patch(bj.id)
        # ... create one job in it, then:
        # await c.JobsApi(cl).delete_job_jobs_id_delete(job_id)
        # now bj is QUEUED with job_ids == []

asyncio.run(main())

Observed batches

On 2026-09-22 five batches reached this state and held all five queue slots for about two hours:

844305   queued 07:18:02 UTC
844307   queued 07:31:01 UTC
844309   queued 07:57:30 UTC
844310   queued 08:09:49 UTC
844311   queued 08:21:28 UTC

They were cleared some time later without any action on our side, so this is not urgent for us -
but the state itself is reachable by any user with two documented calls.

Expected behaviour

Either of these would be enough:

  • a DELETE /batch_jobs/{id} endpoint, or
  • automatic cleanup of a QUEUED batch whose last job is removed, or
  • finish accepting a zero-job batch so the state is at least terminable from the client.

Nice to have: have the 429 name the batches occupying the slots, and/or expose a queue position.

Why the batch was never scheduled (related, and our mistake, not a platform bug)

Worth recording because it is easy to hit and hard to diagnose. The wait argument counts cycles
of 20 ns, and per the operational-specifics documentation the idle time is paid once per shot, so a
circuit's execution time scales with shots x delay. job_execution_time_limit (300 s for Tuna-17)
applies to the whole batch. A single circuit using wait(65536) needs roughly 430 s at 65536 shots,
so a batch containing it can never fit the budget - it just never runs. One circuit per batch, and
keeping the delay within the budget, resolves it.

Dominant language
Python
Stars
0
Forks
1
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from QuTech-Delft/compute-api-client

All issues in QuTech-Delft/compute-api-client

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.