Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[feature] Granular timeout management

Open
#11 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Active
Tech stack
github-actions, python
Domain
ci-cd, testing

Research direction

Start by reviewing the GitHub Actions job steps and the linked debug-timeout-minutes-steps-gha-test experiment. Compare tox's --exit-and-dump-after and possible --fail-fast behavior with step-level timeout-minutes, along with the mentioned pytest, ansible-test, CPython, and faulthandler options. Done means proposing a clear timeout interface and identifying how callers can override defaults.

Written by the indexing model from the issue text.

Description

The current timeout strategy is bound to the overall job duration, which is problematic in cases when we can know early that it is doomed to fail but still proceed with running it.

For example, when a job is cancelled because another matrix job fails, it goes into the rerun mode which buys it 5 minutes of overtime (due to how GH handles cancellations).

Another example is when third party dependencies like PPA are flaky or down. The test dep installation can make the overall job duration longer, thus increasing a chance of hitting the timeout and having the GHA platform killing the job close to its completion. Even if it would've succeeded otherwise.

So we likely need to integrate tox's --exit-and-dump-after (and possibly --fail-fast). And maybe encourage the use of similar features in pytest. ansible-test and CPython's test runners also have features of controlled timeouts. Plus there's faulthandler that can dump current traces, worth mentioning.

Back to GHA's timeouts. I think, we should explore step-level timeouts. We would be able to catch certain parts of jobs getting out of line more predictably, but it's unclear what the interface could be for the callers to override the defaults.

I've tested in https://github.com/webknjaz/debug-timeout-minutes-steps-gha-test/actions/runs/37228799583/job/111513907748 that it's possible to set timeout-minutes for steps dynamically, computed earlier in the job as follows:

jobs:
  build:
    runs-on: ubuntu-24.04-arm
    steps:
    - name: Compute step timeouts
      id: timer-values
      run: |-
        echo val1=1 >> "${GITHUB_OUTPUT}"
        echo val2=2 >> "${GITHUB_OUTPUT}"
    - name: SHOULD stop after a minute
      if: always()
      run: sleep 300
      timeout-minutes: ${{ fromJSON(steps.timer-values.outputs.val1) }}
    - name: SHOULD stop after two minutes
      if: always()
      run: sleep 300
      timeout-minutes: ${{ fromJSON(steps.timer-values.outputs.val2) }}
Dominant language
No language data
Stars
6
Forks
2
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from tox-dev/workflow

All issues in tox-dev/workflow

Similar issues

More DevOps issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.