Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[v3] FrameworkProcessor and ModelTrainer: 4 regressions (including dropping CodeArtifact support) from v2 migration

Đang mở
#5,765 2 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

  • #5790 của @aviruthen — đã đóng, không merge

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
aws, python, pytorch

Hướng nghiên cứu

Bắt đầu với sagemaker/core/resources.py để xử lý session, sagemaker-core/src/sagemaker/core/processing.py cho _package_code và _generate_framework_script, và sagemaker.train.templates cho INSTALL_REQUIREMENTS. So sánh hành vi CodeArtifact của v2 trong PR #4145 với hành vi của training-toolkit. Hoàn tất nghĩa là cả bốn đường dẫn hồi quy đều giữ nguyên session được cung cấp, tuân theo code_location và cài đặt các requirements riêng tư thông qua CodeArtifact.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Additional context

The following four issues were found during a migration from sagemaker==2.257.1 to sagemaker==3.7.1, moving training jobs from sagemaker.pytorch.PyTorch to sagemaker.train.ModelTrainer and processing jobs from sagemaker.pytorch.processing.PyTorchProcessor to sagemaker.core.processing.FrameworkProcessor.


Describe the bug

Four regressions found when migrating from v2 to v3, specifically around FrameworkProcessor (processing jobs) and ModelTrainer (training jobs). All four worked correctly in v2.

System information

  • SageMaker Python SDK version: 3.7.1
  • Framework name: PyTorch
  • Framework version: 2.10
  • Python version: 3.13
  • CPU or GPU: Both
  • Custom Docker image: N

Bug 1: wait=True does not respect sagemaker session

Affects: ModelTrainer.train(wait=True) and FrameworkProcessor.run(wait=True)

ProcessingJob.refresh() and TrainingJob.refresh() use Base.get_sagemaker_client() — a global/default client — instead of the sagemaker_session passed to the processor/trainer. This fails with NoCredentialsError when using assumed-role sessions (via STS).

To reproduce:

import boto3
from sagemaker.core.helper.session_helper import Session
from sagemaker.core.processing import FrameworkProcessor
from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

# Assumed-role session
sts = boto3.client("sts")
assumed = sts.assume_role(RoleArn="arn:aws:iam::123456789:role/MyRole", RoleSessionName="test")
creds = assumed["Credentials"]
assumed_session = boto3.Session(
    aws_access_key_id=creds["AccessKeyId"],
    aws_secret_access_key=creds["SecretAccessKey"],
    aws_session_token=creds["SessionToken"],
    region_name="us-west-2",
)
sm_session = Session(boto_session=assumed_session)

# FrameworkProcessor — job created OK, wait fails
processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    sagemaker_session=sm_session,
)
processor.run(code="my_script.py", source_dir="src", wait=True)
# → NoCredentialsError

# ModelTrainer — same issue
trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(entry_script="train.py", source_dir="src"),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    sagemaker_session=sm_session,
)
trainer.train(wait=True)
# → NoCredentialsError

Root cause (sagemaker/core/resources.py):

# ProcessingJob.refresh() / TrainingJob.refresh()
client = Base.get_sagemaker_client()  # ← ignores the session
response = client.describe_processing_job(**operation_input_args)

v2 behaviour — ProcessingJob.wait() used the session directly:

def wait(self, logs=True):
    if logs:
        self.sagemaker_session.logs_for_processing_job(self.job_name, wait=True)
    else:
        self.sagemaker_session.wait_for_processing_job(self.job_name)

Bug 2: FrameworkProcessor.code_location is accepted but ignored

FrameworkProcessor.__init__ accepts code_location and stores it as self.code_location. The docstring states it controls where code is uploaded. However, _package_code ignores it and always uploads to self.sagemaker_session.default_bucket().

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    code_location="s3://my-custom-bucket",  # ← ignored
)
processor.run(code="my_script.py", source_dir="src", wait=False)
# Code uploads to s3://sagemaker-us-west-2-123456789/... instead of s3://my-custom-bucket/...

Root cause (sagemaker-core/src/sagemaker/core/processing.py, _package_code):

s3_uri = s3.s3_path_join(
    "s3://",
    self.sagemaker_session.default_bucket(),  # ← always uses default bucket
    self.sagemaker_session.default_bucket_prefix or "",
    job_name, "source", "sourcedir.tar.gz",
)

self.code_location is never referenced.

v2 behaviour — FrameworkProcessor delegated to an estimator that honored code_location:

# v2 FrameworkProcessor._create_estimator
return self.estimator_cls(
    ...
    code_location=self.code_location,
    ...
)

Note: ModelTrainer does not offer code_location at all — it always uses session.default_bucket(). Suggested fix: either remove code_location from FrameworkProcessor to align with ModelTrainer, or update _package_code to use it when set.


Bugs 3 and 4 are both about losing the ability to install requirements.txt dependencies from a CodeArtifact repository.

  1. PyTorch training and inference containers supported this via the CA_REPOSITORY_ARN environment variable (see https://github.com/aws/deep-learning-containers/issues/2509 for details)
  2. https://github.com/aws/sagemaker-python-sdk/pull/4145 extended that support to processing jobs by exposing codeartifact_repo_arn on FrameworkProcessor.run()

For context, the new ray-based inference containers are also adding codeartifact support via CA_REPOSITORY_ARN environment variable as can be seen in https://github.com/aws/deep-learning-containers/blob/0fc07f317a4db68ff728274070fbe332dde1ca26/scripts/ray/sagemaker_serve.py#L68


Bug 3: CodeArtifact support missing from FrameworkProcessor

In v2, FrameworkProcessor.run() accepted codeartifact_repo_arn (added in PR #4145). This configured pip inside the container to authenticate with CodeArtifact before installing requirements.txt.

In v3, FrameworkProcessor.run() does not accept codeartifact_repo_arn, and _generate_framework_script has no CodeArtifact support. The generated runproc.sh runs pip install -r requirements.txt without authentication, failing for packages hosted on private CodeArtifact repositories.

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
)

# v2 supported: processor.run(..., codeartifact_repo_arn="arn:aws:codeartifact:us-west-2:123:repository/domain/repo")
# v3 does not — parameter doesn't exist

processor.run(code="my_script.py", source_dir="src", wait=False)
# Container fails: pip install -r requirements.txt → package not found (private CodeArtifact repo)

v2 behaviour — _generate_framework_script injected CodeArtifact login into runproc.sh:

if [[ -f 'requirements.txt' ]]; then
    if ! hash aws 2>/dev/null; then
        echo "AWS CLI is not installed. Skipping CodeArtifact login."
    else
        aws codeartifact login --tool pip --domain {domain} --domain-owner {owner} --repository {repository} --region {region}
    fi
    pip install -r requirements.txt
fi

Suggested fix: Port codeartifact_repo_arn and _get_codeartifact_command from v2's PR #4145 into v3's FrameworkProcessor.


Bug 4: ModelTrainer bypasses sagemaker-training-toolkit, losing CodeArtifact support for requirements.txt

Affects: ModelTrainer.train() with SourceCode(requirements="requirements.txt")

ModelTrainer overrides the container's ENTRYPOINT with its own sm_train.sh driver script. This bypasses the sagemaker-training-toolkit installed in the container, which handled requirements.txt installation with CodeArtifact support (via the CA_REPOSITORY_ARN environment variable, added in sagemaker-training-toolkit#187).

The v3 sm_train.sh template for requirements installation is a bare pip install:

# from sagemaker.train.templates.INSTALL_REQUIREMENTS
echo "Installing requirements"
$SM_PIP_CMD install -r {requirements_file}

It does not check for CA_REPOSITORY_ARN or configure pip to use CodeArtifact. The env var is passed to the container but nothing reads it.

To reproduce:

from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(
        entry_script="train.py",
        source_dir="src",
        requirements="requirements.txt",  # ← installed without CodeArtifact
    ),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    environment={"CA_REPOSITORY_ARN": "arn:aws:codeartifact:us-west-2:123:repository/domain/repo"},
)
trainer.train(wait=False)
# Container runs: pip install -r requirements.txt (using public PyPI, not CodeArtifact)
# Fails in VPC-isolated environments where PyPI is unreachable

v2 behaviour — the PyTorch estimator used the container's native entrypoint, which invoked sagemaker-training-toolkit. The toolkit checked CA_REPOSITORY_ARN, ran aws codeartifact login --tool pip, then installed requirements. This was added in sagemaker-training-toolkit v4.7.0 and tracked in deep-learning-containers#2509.

Suggested fix: The INSTALL_REQUIREMENTS template in sagemaker.train.templates should check for CA_REPOSITORY_ARN and configure pip accordingly before installing, matching the behaviour of sagemaker-training-toolkit.

Ngôn ngữ chính
Python
Star
2.3k
Fork
1.3k
Merge trung bình
3 ngày 7 giờ
Pull request đã merge (30 ngày)
87

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của aws/sagemaker-python-sdk

Tất cả issue của aws/sagemaker-python-sdk

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.