Training job unit tests can outlive their service mocks and hang teardown
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 72/100
Research direction
Start with test_run_called_twice_raises in the AutoML forecasting, image, tabular, text and video unit test files, and with test_run_returns_none_if_no_model_to_upload in the custom training job tests. In each async case, join the first job with job.wait() after the duplicate-run assertion, as the sync-free custom-training tests already do. Run the diagnostic script from the issue. Done when it passes for all 22 cases and teardown no longer waits on real pipeline-client retries.
Written by the indexing model from the issue text.
Description
The async cases of test_run_called_twice_raises in the AutoML forecasting, image, tabular, text and video modules can return while the first training job is still running. The second run() correctly raises RuntimeError, but the test does not wait for the first job before pytest removes the service mocks.
The same lifetime issue also occurs in test_run_returns_none_if_no_model_to_upload for CustomTrainingJob, CustomContainerTrainingJob and CustomPythonPackageTrainingJob. Those tests assert that run() returns None, then return without joining the job.
During a bounded investigation, a forecasting worker entered the real PipelineServiceClient.get_training_pipeline after those mocks were removed, while class teardown waited in initializer.global_pool.shutdown(wait=True). This blocked the unit-test run on RPC retries in an offline environment. A subsequent parallel run also captured a real pipeline-client retry during the custom-container no-model test, whose teardown took 80.49 seconds despite the test ultimately being reported as passed.
Environment
- Current main:
7f08bd6c1dbb99f8a00922ae13c488bda3ece7b9. - Debian 12 container, Linux/WSL2, CPython 3.14.7, pip 26.2.1.
- SDK imported from the checkout (
2.4.0), with the standard unit-test dependencies. No cloud credentials or external network. - The existing environment's
pip checkreports invalid MLflow optional-extra metadata (scikit-learn >=1.0.*) on both unchanged and modified main;uv pip checkpasses. No dependency versions were changed for this investigation.
Reproduction
The ordinary hang depends on worker scheduling. The diagnostic below makes the lifetime error deterministic for the existing SequenceToSequencePlusForecastingTrainingJob-False case. It holds the worker just before model lookup, preserves the original duplicate-run assertion, and checks for unfinished work when the test returns. It then releases and joins the worker while the service mocks are still active, so the diagnostic fails rather than hanging in teardown.
Save the script outside the checkout as /tmp/automl-lifetime-repro.py, then run from the repository root in its unit-test environment:
python -I -B /tmp/automl-lifetime-repro.py
Expected: background work completes before the test returns. On the main commit above, the script fails with AssertionError: Test returned with background work pending. Adding an async job.wait() after the existing exception assertion makes it pass. The same barrier check across the five AutoML modules and the three custom no-model methods found 11 failing async cases and 11 passing sync cases; all 22 pass with the final cleanup, and removing the cleanup restores the same 11 failures.
Standalone diagnostic (verified on unchanged main and with the cleanup)
"""Run from the repository root with its unit-test dependencies installed."""
import inspect
from pathlib import Path
import sys
import threading
from unittest import mock
import pytest
sys.path.insert(0, str(Path.cwd()))
from google.cloud.aiplatform import base, training_jobs
class LifetimeCheck:
@pytest.hookimpl(hookwrapper=True, tryfirst=True)
def pytest_runtest_call(self, item):
entered, release = threading.Event(), threading.Event()
jobs = []
original_get = training_jobs._TrainingJob._get_model
original_wait = base.FutureManager.wait
def held_get(job, *args, **kwargs):
if threading.current_thread() is not threading.main_thread():
jobs.append(job)
entered.set()
assert release.wait(10), "Worker barrier expired"
return original_get(job, *args, **kwargs)
def explicit_wait(job):
release.set()
return original_wait(job)
with mock.patch.object(training_jobs._TrainingJob, "_get_model", held_get), mock.patch.object(base.FutureManager, "wait", explicit_wait):
outcome = yield
observed = entered.wait(5)
completed = observed and all(job._are_futures_done() for job in jobs)
# Safely finish the worker while the test's service mocks still exist.
release.set()
for job in jobs:
original_wait(job)
outcome.get_result()
assert observed, "Worker did not reach the model boundary"
assert completed, "Test returned with background work pending"
print("SDK source:", inspect.getfile(training_jobs), flush=True)
raise SystemExit(pytest.main([
"-q", "-p", "no:warnings", "-p", "no:cacheprovider",
"tests/unit/aiplatform/test_automl_forecasting_training_jobs.py::"
"TestForecastingTrainingJob::test_run_called_twice_raises"
"[SequenceToSequencePlusForecastingTrainingJob-False]",
], plugins=[LifetimeCheck()]))
I have a local test-only cleanup that adds if not sync: job.wait() after the existing assertion in these eight methods across six files. The three equivalent custom-training duplicate-run tests already use that pattern. Would this be an appropriate focused test-cleanup contribution?
- Dominant language
- Python
- Stars
- 907
- Forks
- 469
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 38
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from googleapis/python-aiplatform
-
Protobuf 7.35.1+ supportPossibly taken @garethknowles claimed this 30 days ago. Openapi: vertex-ai
Difficulty 1/5 Under an hour Newbie friendliness 85/100
googleapis/python-aiplatform#7132 · 1 reaction ·
Maintainers usually reply within 2 days
-
AgentEngineConfig.agent_framework Literal rejects "a2a" — out of sync with _SUPPORTED_AGENT_FRAMEWORKS and the GA A2A templateMay be free again A pull request for this issue was closed without being merged. Openapi: vertex-ai
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
googleapis/python-aiplatform#7097 ·
Maintainers usually reply within 2 days
-
CustomContainerTrainingJob.run drops max_wait_duration=0 instead of requesting indefinite DWS waitPossibly taken @hugosmoreira claimed this 58 days ago. Openapi: vertex-ai
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
googleapis/python-aiplatform#7067 · 1 comment ·
Maintainers usually reply within 2 days
-
`A2aAgent` builds agent-card `url` from default region (us-central1), ignoring `GOOGLE_CLOUD_AGENT_ENGINE_LOCATION`Possibly taken @sushicw claimed this 121 days ago. Openapi: vertex-ai
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
googleapis/python-aiplatform#6877 ·
Maintainers usually reply within 2 days
-
api: vertex-ai
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
googleapis/python-aiplatform#6865 · 1 comment ·
Maintainers usually reply within 2 days
All issues in googleapis/python-aiplatform
Similar issues
-
namespace operations
Difficulty 1/5 Under an hour Newbie friendliness 72/100
EclipseFdn/open-vsx.org#14043 ·
Maintainers usually reply within 1 day
-
feedback simulation workshop
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
githubnext/gh-aw-workshop#4455 ·
Maintainers usually reply within 1 day
-
Triage 🩺
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOpenneeds-triage
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
krkn-chaos/krkn#1627 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NousResearch/hermes-agent#136483 ·
Maintainers usually reply within 1 day