Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Idle CU sweep cannot reclaim a computing unit whose execution is stuck in a non-terminal state

Open
#8,618 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
38/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
kubernetes, scala

Research direction

Start with ComputingUnitHelpers.reconcileVanishedKubernetesUnits and the idle Kubernetes computing unit sweep, then trace how workflow_executions status and last_update_time are used. Compare the existing listing-triggered reconciliation with the proposed scheduled pod-state check. Done means a live pod with a stale non-terminal execution no longer keeps an idle computing unit unreclaimed.

Written by the indexing model from the issue text.

Description

Feature Summary

Follow-up to #6046.

Problem

The idle Kubernetes computing unit sweep treats a computing unit as busy whenever any of its
workflow_executions rows carries a non-terminal status code (status NOT IN (3, 4, 5)). That
test has no time bound, so an execution row left stuck in a non-terminal state keeps its computing unit off the sweep indefinitely, even though the unit is doing no work.

But the problem is, nothing else reclaims it either: ComputingUnitHelpers.reconcileVanishedKubernetesUnits only
runs when someone calls a listing endpoint, and it only checks whether the pod is already gone,
not whether the execution row is stuck somewhere. A live pod with a stuck row is missed on both paths.

maptoStatusCode also maps UNKNOWN (and TERMINATED) to -1, which is not in {3,4,5}. A row whose final status is -1 pins its unit off the sweep permanently. Fixing that means changing the codes amber writes, not this sweep's predicate.

Proposed Solution or Design
  • Ignore a non-terminal execution row whose last_update_time is older than its own timeout.
  • Check execution status codes against actual pod state on a schedule, rather than only on a
    listing request.

Either is a larger change than #6046 should carry, hence this follow-up.

Affected Area

No response

Dominant language
Scala
Stars
316
Forks
192
Avg merge
4d 17h
Merged PRs (30d)
141

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/texera

All issues in apache/texera

Similar issues

More Scala issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.