ExponentialBackoff overflows timedelta above attempt 46

Open Beginner friendly
#47 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Start at threadmill.retry.ExponentialBackoff.call and inspect how its delay is calculated before min() is applied. Reproduce the issue with the provided attempt loop, then check Executor.retry_delay to confirm the callback no longer fails at high attempts. Done means attempts through max_retries remain capped at max_delay instead of ending at the overflow.

Written by the indexing model from the issue text.

Description

ExponentialBackoff.__call__ computes

delay = min(self.base_delay * (self.factor**context.attempt), self.max_delay)

so the product is evaluated before min() clamps it. With the default shape (base delay 1 s, factor 2.0, max delay 1 h) the product exceeds the timedelta limit at attempt 47:

OverflowError: days=1628906115; must have magnitude <= 999999999
Repro (threadmill 0.7.1)
import datetime
from types import SimpleNamespace
from threadmill.retry import ExponentialBackoff

policy = ExponentialBackoff(
    base_delay=datetime.timedelta(seconds=1),
    max_delay=datetime.timedelta(hours=1),
    factor=2.0,
    max_retries=720,
)

def context(attempt):
    error = SimpleNamespace(exception_class=ValueError)
    return SimpleNamespace(attempt=attempt, task_result=SimpleNamespace(errors=[error]))

for attempt in range(1, 60):
    try:
        print(attempt, policy(context(attempt)))
    except Exception as exc:
        print(attempt, type(exc).__name__, exc)
        break

Attempts 1 to 46 return a delay (2 s doubling to the 1 h cap at attempt 12), attempt 47 raises.

Why it matters

Executor.retry_delay catches the exception, logs Retry callback failed, and returns None, so the backend acknowledges the result and the retry chain ends. The task looks like it exhausted its policy, but max_retries was never reached: a policy of 720 attempts really stops after 46 retries.

Suggested fix

Clamp in seconds before building the timedelta, and derive the capped attempt count from max_delay so the exponential is never evaluated past the cap:

seconds = min(
    self.base_delay.total_seconds() * self.factor**context.attempt,
    self.max_delay.total_seconds(),
)
return datetime.timedelta(seconds=seconds)

A plain min() on the two timedelta values does not help on its own, because the product still overflows before the comparison.

Found while bounding the spam scan retry budget in codingjoe/relay#230.

Dominant language
Python
Stars
12
Forks
1
Avg merge
1d 1h
Merged PRs (30d)
10

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from codingjoe/threadmill

All issues in codingjoe/threadmill

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.