CustomContainerTrainingJob.run drops max_wait_duration=0 instead of requesting indefinite DWS wait

Open Beginner friendly
#7,067 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Quiet
Tech stack
python

Research direction

Start in google/cloud/aiplatform/training_jobs.py at _prepare_training_task_inputs_and_output_dir, then compare the lower-level CustomJob implementation referenced in the issue. Verify the scheduling inputs for omitted, zero, and positive max_wait_duration values; done means zero is serialized as 0s while omission remains distinct and positive durations are preserved.

Written by the indexing model from the issue text.

Description

api: vertex-ai

Thanks for maintaining this library. CustomContainerTrainingJob.run(max_wait_duration=0) drops the value because the high-level training-job implementation uses a truthiness check. This prevents callers from requesting the documented indefinite wait for DWS Flex Start.

Environment details
  • OS type and version: macOS
  • Python version: 3.12
  • google-cloud-aiplatform versions: reproduced in 1.148.1; source inspection confirms the same behavior in 1.150.0, 1.163.0, and current main
Steps to reproduce
  1. Create a CustomContainerTrainingJob using any valid project, staging bucket, and container image.
  2. Run it with Flex Start and max_wait_duration=0.
  3. Inspect the scheduling inputs produced by _prepare_training_task_inputs_and_output_dir or the resulting TrainingPipeline request.
Code example
job.run(
    scheduling_strategy=custom_job.Scheduling.Strategy.FLEX_START,
    max_wait_duration=0,
)

Observed scheduling value:

max_wait_duration: None

Expected scheduling value:

max_wait_duration: 0s

For comparison, a positive input such as 14400 is preserved as 14400s.

Cause

The high-level training-job path converts the duration using a truthiness check, so zero follows the same branch as None:

https://github.com/googleapis/python-aiplatform/blob/v1.163.0/google/cloud/aiplatform/training_jobs.py#L1663-L1667

The generated API contract says that explicit zero means indefinite waiting, while omission defaults to 24 hours:

https://github.com/googleapis/python-aiplatform/blob/v1.163.0/google/cloud/aiplatform_v1/types/custom_job.py#L577-L582

max_wait_duration is a presence-aware google.protobuf.Duration, so explicit zero and an omitted field are distinct requests.

A previous fix correctly changed the lower-level CustomJob and hyperparameter-tuning paths to use is not None, explicitly noting that 0 is valid:

https://github.com/googleapis/python-aiplatform/commit/d9675fdf051233539f478187143f2833fd6e6af0

That commit did not update the CustomContainerTrainingJob path in training_jobs.py.

Suggested fix

Use an explicit is not None check in _prepare_training_task_inputs_and_output_dir, consistent with the lower-level implementation, and add tests covering omitted, zero, and positive durations.

Stack trace

No exception is raised. The problem is a silent request-semantics change: 0 is serialized as omission.

Dominant language
Python
Stars
907
Forks
467
Avg merge
1d 13h
Merged PRs (30d)
44

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from googleapis/python-aiplatform

All issues in googleapis/python-aiplatform

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.