Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Training job unit tests can outlive their service mocks and hang teardown

Open Beginner friendly
#7,208 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
72/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
gcp, python
Domain
testing

Research direction

Start with test_run_called_twice_raises in the AutoML forecasting, image, tabular, text and video unit test files, and with test_run_returns_none_if_no_model_to_upload in the custom training job tests. In each async case, join the first job with job.wait() after the duplicate-run assertion, as the sync-free custom-training tests already do. Run the diagnostic script from the issue. Done when it passes for all 22 cases and teardown no longer waits on real pipeline-client retries.

Written by the indexing model from the issue text.

Description

api: vertex-ai

The async cases of test_run_called_twice_raises in the AutoML forecasting, image, tabular, text and video modules can return while the first training job is still running. The second run() correctly raises RuntimeError, but the test does not wait for the first job before pytest removes the service mocks.

The same lifetime issue also occurs in test_run_returns_none_if_no_model_to_upload for CustomTrainingJob, CustomContainerTrainingJob and CustomPythonPackageTrainingJob. Those tests assert that run() returns None, then return without joining the job.

During a bounded investigation, a forecasting worker entered the real PipelineServiceClient.get_training_pipeline after those mocks were removed, while class teardown waited in initializer.global_pool.shutdown(wait=True). This blocked the unit-test run on RPC retries in an offline environment. A subsequent parallel run also captured a real pipeline-client retry during the custom-container no-model test, whose teardown took 80.49 seconds despite the test ultimately being reported as passed.

Environment
  • Current main: 7f08bd6c1dbb99f8a00922ae13c488bda3ece7b9.
  • Debian 12 container, Linux/WSL2, CPython 3.14.7, pip 26.2.1.
  • SDK imported from the checkout (2.4.0), with the standard unit-test dependencies. No cloud credentials or external network.
  • The existing environment's pip check reports invalid MLflow optional-extra metadata (scikit-learn >=1.0.*) on both unchanged and modified main; uv pip check passes. No dependency versions were changed for this investigation.
Reproduction

The ordinary hang depends on worker scheduling. The diagnostic below makes the lifetime error deterministic for the existing SequenceToSequencePlusForecastingTrainingJob-False case. It holds the worker just before model lookup, preserves the original duplicate-run assertion, and checks for unfinished work when the test returns. It then releases and joins the worker while the service mocks are still active, so the diagnostic fails rather than hanging in teardown.

Save the script outside the checkout as /tmp/automl-lifetime-repro.py, then run from the repository root in its unit-test environment:

python -I -B /tmp/automl-lifetime-repro.py

Expected: background work completes before the test returns. On the main commit above, the script fails with AssertionError: Test returned with background work pending. Adding an async job.wait() after the existing exception assertion makes it pass. The same barrier check across the five AutoML modules and the three custom no-model methods found 11 failing async cases and 11 passing sync cases; all 22 pass with the final cleanup, and removing the cleanup restores the same 11 failures.

Standalone diagnostic (verified on unchanged main and with the cleanup)
"""Run from the repository root with its unit-test dependencies installed."""
import inspect
from pathlib import Path
import sys
import threading
from unittest import mock

import pytest

sys.path.insert(0, str(Path.cwd()))
from google.cloud.aiplatform import base, training_jobs


class LifetimeCheck:
    @pytest.hookimpl(hookwrapper=True, tryfirst=True)
    def pytest_runtest_call(self, item):
        entered, release = threading.Event(), threading.Event()
        jobs = []
        original_get = training_jobs._TrainingJob._get_model
        original_wait = base.FutureManager.wait

        def held_get(job, *args, **kwargs):
            if threading.current_thread() is not threading.main_thread():
                jobs.append(job)
                entered.set()
                assert release.wait(10), "Worker barrier expired"
            return original_get(job, *args, **kwargs)

        def explicit_wait(job):
            release.set()
            return original_wait(job)

        with mock.patch.object(training_jobs._TrainingJob, "_get_model", held_get), mock.patch.object(base.FutureManager, "wait", explicit_wait):
            outcome = yield
            observed = entered.wait(5)
            completed = observed and all(job._are_futures_done() for job in jobs)
            # Safely finish the worker while the test's service mocks still exist.
            release.set()
            for job in jobs:
                original_wait(job)
            outcome.get_result()
            assert observed, "Worker did not reach the model boundary"
            assert completed, "Test returned with background work pending"


print("SDK source:", inspect.getfile(training_jobs), flush=True)
raise SystemExit(pytest.main([
    "-q", "-p", "no:warnings", "-p", "no:cacheprovider",
    "tests/unit/aiplatform/test_automl_forecasting_training_jobs.py::"
    "TestForecastingTrainingJob::test_run_called_twice_raises"
    "[SequenceToSequencePlusForecastingTrainingJob-False]",
], plugins=[LifetimeCheck()]))

I have a local test-only cleanup that adds if not sync: job.wait() after the existing assertion in these eight methods across six files. The three equivalent custom-training duplicate-run tests already use that pattern. Would this be an appropriate focused test-cleanup contribution?

Dominant language
Python
Stars
907
Forks
469
Avg merge
1d 15h
Merged PRs (30d)
38

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from googleapis/python-aiplatform

All issues in googleapis/python-aiplatform

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.