Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[v3] FrameworkProcessor and ModelTrainer: 4 regressions (including dropping CodeArtifact support) from v2 migration

未关闭
#5,765 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 2 天内回复

还没有人认领这个 Issue。

  • #5790 来自 @aviruthen —— 已关闭,未合并

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
aws, python, pytorch

调研方向

从 sagemaker/core/resources.py 开始了解会话处理,从 sagemaker-core/src/sagemaker/core/processing.py 查看 _package_code 和 _generate_framework_script,并从 sagemaker.train.templates 查看 INSTALL_REQUIREMENTS。比较 PR #4145 中 v2 的 CodeArtifact 行为与 training-toolkit 的行为。完成的标准是,四条回归路径都保留提供的会话、遵循 code_location,并通过 CodeArtifact 安装私有 requirements。

由索引模型根据 Issue 内容生成。

描述

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Additional context

The following four issues were found during a migration from sagemaker==2.257.1 to sagemaker==3.7.1, moving training jobs from sagemaker.pytorch.PyTorch to sagemaker.train.ModelTrainer and processing jobs from sagemaker.pytorch.processing.PyTorchProcessor to sagemaker.core.processing.FrameworkProcessor.


Describe the bug

Four regressions found when migrating from v2 to v3, specifically around FrameworkProcessor (processing jobs) and ModelTrainer (training jobs). All four worked correctly in v2.

System information

  • SageMaker Python SDK version: 3.7.1
  • Framework name: PyTorch
  • Framework version: 2.10
  • Python version: 3.13
  • CPU or GPU: Both
  • Custom Docker image: N

Bug 1: wait=True does not respect sagemaker session

Affects: ModelTrainer.train(wait=True) and FrameworkProcessor.run(wait=True)

ProcessingJob.refresh() and TrainingJob.refresh() use Base.get_sagemaker_client() — a global/default client — instead of the sagemaker_session passed to the processor/trainer. This fails with NoCredentialsError when using assumed-role sessions (via STS).

To reproduce:

import boto3
from sagemaker.core.helper.session_helper import Session
from sagemaker.core.processing import FrameworkProcessor
from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

# Assumed-role session
sts = boto3.client("sts")
assumed = sts.assume_role(RoleArn="arn:aws:iam::123456789:role/MyRole", RoleSessionName="test")
creds = assumed["Credentials"]
assumed_session = boto3.Session(
    aws_access_key_id=creds["AccessKeyId"],
    aws_secret_access_key=creds["SecretAccessKey"],
    aws_session_token=creds["SessionToken"],
    region_name="us-west-2",
)
sm_session = Session(boto_session=assumed_session)

# FrameworkProcessor — job created OK, wait fails
processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    sagemaker_session=sm_session,
)
processor.run(code="my_script.py", source_dir="src", wait=True)
# → NoCredentialsError

# ModelTrainer — same issue
trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(entry_script="train.py", source_dir="src"),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    sagemaker_session=sm_session,
)
trainer.train(wait=True)
# → NoCredentialsError

Root cause (sagemaker/core/resources.py):

# ProcessingJob.refresh() / TrainingJob.refresh()
client = Base.get_sagemaker_client()  # ← ignores the session
response = client.describe_processing_job(**operation_input_args)

v2 behaviour — ProcessingJob.wait() used the session directly:

def wait(self, logs=True):
    if logs:
        self.sagemaker_session.logs_for_processing_job(self.job_name, wait=True)
    else:
        self.sagemaker_session.wait_for_processing_job(self.job_name)

Bug 2: FrameworkProcessor.code_location is accepted but ignored

FrameworkProcessor.__init__ accepts code_location and stores it as self.code_location. The docstring states it controls where code is uploaded. However, _package_code ignores it and always uploads to self.sagemaker_session.default_bucket().

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    code_location="s3://my-custom-bucket",  # ← ignored
)
processor.run(code="my_script.py", source_dir="src", wait=False)
# Code uploads to s3://sagemaker-us-west-2-123456789/... instead of s3://my-custom-bucket/...

Root cause (sagemaker-core/src/sagemaker/core/processing.py, _package_code):

s3_uri = s3.s3_path_join(
    "s3://",
    self.sagemaker_session.default_bucket(),  # ← always uses default bucket
    self.sagemaker_session.default_bucket_prefix or "",
    job_name, "source", "sourcedir.tar.gz",
)

self.code_location is never referenced.

v2 behaviour — FrameworkProcessor delegated to an estimator that honored code_location:

# v2 FrameworkProcessor._create_estimator
return self.estimator_cls(
    ...
    code_location=self.code_location,
    ...
)

Note: ModelTrainer does not offer code_location at all — it always uses session.default_bucket(). Suggested fix: either remove code_location from FrameworkProcessor to align with ModelTrainer, or update _package_code to use it when set.


Bugs 3 and 4 are both about losing the ability to install requirements.txt dependencies from a CodeArtifact repository.

  1. PyTorch training and inference containers supported this via the CA_REPOSITORY_ARN environment variable (see https://github.com/aws/deep-learning-containers/issues/2509 for details)
  2. https://github.com/aws/sagemaker-python-sdk/pull/4145 extended that support to processing jobs by exposing codeartifact_repo_arn on FrameworkProcessor.run()

For context, the new ray-based inference containers are also adding codeartifact support via CA_REPOSITORY_ARN environment variable as can be seen in https://github.com/aws/deep-learning-containers/blob/0fc07f317a4db68ff728274070fbe332dde1ca26/scripts/ray/sagemaker_serve.py#L68


Bug 3: CodeArtifact support missing from FrameworkProcessor

In v2, FrameworkProcessor.run() accepted codeartifact_repo_arn (added in PR #4145). This configured pip inside the container to authenticate with CodeArtifact before installing requirements.txt.

In v3, FrameworkProcessor.run() does not accept codeartifact_repo_arn, and _generate_framework_script has no CodeArtifact support. The generated runproc.sh runs pip install -r requirements.txt without authentication, failing for packages hosted on private CodeArtifact repositories.

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
)

# v2 supported: processor.run(..., codeartifact_repo_arn="arn:aws:codeartifact:us-west-2:123:repository/domain/repo")
# v3 does not — parameter doesn't exist

processor.run(code="my_script.py", source_dir="src", wait=False)
# Container fails: pip install -r requirements.txt → package not found (private CodeArtifact repo)

v2 behaviour — _generate_framework_script injected CodeArtifact login into runproc.sh:

if [[ -f 'requirements.txt' ]]; then
    if ! hash aws 2>/dev/null; then
        echo "AWS CLI is not installed. Skipping CodeArtifact login."
    else
        aws codeartifact login --tool pip --domain {domain} --domain-owner {owner} --repository {repository} --region {region}
    fi
    pip install -r requirements.txt
fi

Suggested fix: Port codeartifact_repo_arn and _get_codeartifact_command from v2's PR #4145 into v3's FrameworkProcessor.


Bug 4: ModelTrainer bypasses sagemaker-training-toolkit, losing CodeArtifact support for requirements.txt

Affects: ModelTrainer.train() with SourceCode(requirements="requirements.txt")

ModelTrainer overrides the container's ENTRYPOINT with its own sm_train.sh driver script. This bypasses the sagemaker-training-toolkit installed in the container, which handled requirements.txt installation with CodeArtifact support (via the CA_REPOSITORY_ARN environment variable, added in sagemaker-training-toolkit#187).

The v3 sm_train.sh template for requirements installation is a bare pip install:

# from sagemaker.train.templates.INSTALL_REQUIREMENTS
echo "Installing requirements"
$SM_PIP_CMD install -r {requirements_file}

It does not check for CA_REPOSITORY_ARN or configure pip to use CodeArtifact. The env var is passed to the container but nothing reads it.

To reproduce:

from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(
        entry_script="train.py",
        source_dir="src",
        requirements="requirements.txt",  # ← installed without CodeArtifact
    ),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    environment={"CA_REPOSITORY_ARN": "arn:aws:codeartifact:us-west-2:123:repository/domain/repo"},
)
trainer.train(wait=False)
# Container runs: pip install -r requirements.txt (using public PyPI, not CodeArtifact)
# Fails in VPC-isolated environments where PyPI is unreachable

v2 behaviour — the PyTorch estimator used the container's native entrypoint, which invoked sagemaker-training-toolkit. The toolkit checked CA_REPOSITORY_ARN, ran aws codeartifact login --tool pip, then installed requirements. This was added in sagemaker-training-toolkit v4.7.0 and tracked in deep-learning-containers#2509.

Suggested fix: The INSTALL_REQUIREMENTS template in sagemaker.train.templates should check for CA_REPOSITORY_ARN and configure pip accordingly before installing, matching the behaviour of sagemaker-training-toolkit.

主要语言
Python
星标
2.3k
派生
1.3k
平均合并
3 天 10 小时
30 天内合并 PR
82

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

aws/sagemaker-python-sdk 的其他 Issue

查看 aws/sagemaker-python-sdk 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。