[v3] FrameworkProcessor and ModelTrainer: 4 regressions (including dropping CodeArtifact support) from v2 migration
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
- #5790 của @aviruthen — đã đóng, không merge
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Lĩnh vực
- backend-api-design, cloud, machine-learning
Hướng nghiên cứu
Bắt đầu với sagemaker/core/resources.py để xử lý session, sagemaker-core/src/sagemaker/core/processing.py cho _package_code và _generate_framework_script, và sagemaker.train.templates cho INSTALL_REQUIREMENTS. So sánh hành vi CodeArtifact của v2 trong PR #4145 với hành vi của training-toolkit. Hoàn tất nghĩa là cả bốn đường dẫn hồi quy đều giữ nguyên session được cung cấp, tuân theo code_location và cài đặt các requirements riêng tư thông qua CodeArtifact.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
PySDK Version
- PySDK V2 (2.x)
- PySDK V3 (3.x)
Additional context
The following four issues were found during a migration from sagemaker==2.257.1 to sagemaker==3.7.1, moving training jobs from sagemaker.pytorch.PyTorch to sagemaker.train.ModelTrainer and processing jobs from sagemaker.pytorch.processing.PyTorchProcessor to sagemaker.core.processing.FrameworkProcessor.
Describe the bug
Four regressions found when migrating from v2 to v3, specifically around FrameworkProcessor (processing jobs) and ModelTrainer (training jobs). All four worked correctly in v2.
System information
- SageMaker Python SDK version: 3.7.1
- Framework name: PyTorch
- Framework version: 2.10
- Python version: 3.13
- CPU or GPU: Both
- Custom Docker image: N
Bug 1: wait=True does not respect sagemaker session
Affects: ModelTrainer.train(wait=True) and FrameworkProcessor.run(wait=True)
ProcessingJob.refresh() and TrainingJob.refresh() use Base.get_sagemaker_client() — a global/default client — instead of the sagemaker_session passed to the processor/trainer. This fails with NoCredentialsError when using assumed-role sessions (via STS).
To reproduce:
import boto3
from sagemaker.core.helper.session_helper import Session
from sagemaker.core.processing import FrameworkProcessor
from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode
# Assumed-role session
sts = boto3.client("sts")
assumed = sts.assume_role(RoleArn="arn:aws:iam::123456789:role/MyRole", RoleSessionName="test")
creds = assumed["Credentials"]
assumed_session = boto3.Session(
aws_access_key_id=creds["AccessKeyId"],
aws_secret_access_key=creds["SecretAccessKey"],
aws_session_token=creds["SessionToken"],
region_name="us-west-2",
)
sm_session = Session(boto_session=assumed_session)
# FrameworkProcessor — job created OK, wait fails
processor = FrameworkProcessor(
image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
command=["python3"],
role="arn:aws:iam::123456789:role/MyRole",
instance_count=1,
instance_type="ml.m5.xlarge",
sagemaker_session=sm_session,
)
processor.run(code="my_script.py", source_dir="src", wait=True)
# → NoCredentialsError
# ModelTrainer — same issue
trainer = ModelTrainer(
training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
role="arn:aws:iam::123456789:role/MyRole",
source_code=SourceCode(entry_script="train.py", source_dir="src"),
compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
sagemaker_session=sm_session,
)
trainer.train(wait=True)
# → NoCredentialsError
Root cause (sagemaker/core/resources.py):
# ProcessingJob.refresh() / TrainingJob.refresh()
client = Base.get_sagemaker_client() # ← ignores the session
response = client.describe_processing_job(**operation_input_args)
v2 behaviour — ProcessingJob.wait() used the session directly:
def wait(self, logs=True):
if logs:
self.sagemaker_session.logs_for_processing_job(self.job_name, wait=True)
else:
self.sagemaker_session.wait_for_processing_job(self.job_name)
Bug 2: FrameworkProcessor.code_location is accepted but ignored
FrameworkProcessor.__init__ accepts code_location and stores it as self.code_location. The docstring states it controls where code is uploaded. However, _package_code ignores it and always uploads to self.sagemaker_session.default_bucket().
To reproduce:
processor = FrameworkProcessor(
image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
command=["python3"],
role="arn:aws:iam::123456789:role/MyRole",
instance_count=1,
instance_type="ml.m5.xlarge",
code_location="s3://my-custom-bucket", # ← ignored
)
processor.run(code="my_script.py", source_dir="src", wait=False)
# Code uploads to s3://sagemaker-us-west-2-123456789/... instead of s3://my-custom-bucket/...
Root cause (sagemaker-core/src/sagemaker/core/processing.py, _package_code):
s3_uri = s3.s3_path_join(
"s3://",
self.sagemaker_session.default_bucket(), # ← always uses default bucket
self.sagemaker_session.default_bucket_prefix or "",
job_name, "source", "sourcedir.tar.gz",
)
self.code_location is never referenced.
v2 behaviour — FrameworkProcessor delegated to an estimator that honored code_location:
# v2 FrameworkProcessor._create_estimator
return self.estimator_cls(
...
code_location=self.code_location,
...
)
Note: ModelTrainer does not offer code_location at all — it always uses session.default_bucket(). Suggested fix: either remove code_location from FrameworkProcessor to align with ModelTrainer, or update _package_code to use it when set.
Bugs 3 and 4 are both about losing the ability to install requirements.txt dependencies from a CodeArtifact repository.
- PyTorch training and inference containers supported this via the
CA_REPOSITORY_ARNenvironment variable (see https://github.com/aws/deep-learning-containers/issues/2509 for details) - https://github.com/aws/sagemaker-python-sdk/pull/4145 extended that support to processing jobs by exposing
codeartifact_repo_arnonFrameworkProcessor.run()
For context, the new ray-based inference containers are also adding codeartifact support via CA_REPOSITORY_ARN environment variable as can be seen in https://github.com/aws/deep-learning-containers/blob/0fc07f317a4db68ff728274070fbe332dde1ca26/scripts/ray/sagemaker_serve.py#L68
Bug 3: CodeArtifact support missing from FrameworkProcessor
In v2, FrameworkProcessor.run() accepted codeartifact_repo_arn (added in PR #4145). This configured pip inside the container to authenticate with CodeArtifact before installing requirements.txt.
In v3, FrameworkProcessor.run() does not accept codeartifact_repo_arn, and _generate_framework_script has no CodeArtifact support. The generated runproc.sh runs pip install -r requirements.txt without authentication, failing for packages hosted on private CodeArtifact repositories.
To reproduce:
processor = FrameworkProcessor(
image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
command=["python3"],
role="arn:aws:iam::123456789:role/MyRole",
instance_count=1,
instance_type="ml.m5.xlarge",
)
# v2 supported: processor.run(..., codeartifact_repo_arn="arn:aws:codeartifact:us-west-2:123:repository/domain/repo")
# v3 does not — parameter doesn't exist
processor.run(code="my_script.py", source_dir="src", wait=False)
# Container fails: pip install -r requirements.txt → package not found (private CodeArtifact repo)
v2 behaviour — _generate_framework_script injected CodeArtifact login into runproc.sh:
if [[ -f 'requirements.txt' ]]; then
if ! hash aws 2>/dev/null; then
echo "AWS CLI is not installed. Skipping CodeArtifact login."
else
aws codeartifact login --tool pip --domain {domain} --domain-owner {owner} --repository {repository} --region {region}
fi
pip install -r requirements.txt
fi
Suggested fix: Port codeartifact_repo_arn and _get_codeartifact_command from v2's PR #4145 into v3's FrameworkProcessor.
Bug 4: ModelTrainer bypasses sagemaker-training-toolkit, losing CodeArtifact support for requirements.txt
Affects: ModelTrainer.train() with SourceCode(requirements="requirements.txt")
ModelTrainer overrides the container's ENTRYPOINT with its own sm_train.sh driver script. This bypasses the sagemaker-training-toolkit installed in the container, which handled requirements.txt installation with CodeArtifact support (via the CA_REPOSITORY_ARN environment variable, added in sagemaker-training-toolkit#187).
The v3 sm_train.sh template for requirements installation is a bare pip install:
# from sagemaker.train.templates.INSTALL_REQUIREMENTS
echo "Installing requirements"
$SM_PIP_CMD install -r {requirements_file}
It does not check for CA_REPOSITORY_ARN or configure pip to use CodeArtifact. The env var is passed to the container but nothing reads it.
To reproduce:
from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode
trainer = ModelTrainer(
training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
role="arn:aws:iam::123456789:role/MyRole",
source_code=SourceCode(
entry_script="train.py",
source_dir="src",
requirements="requirements.txt", # ← installed without CodeArtifact
),
compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
environment={"CA_REPOSITORY_ARN": "arn:aws:codeartifact:us-west-2:123:repository/domain/repo"},
)
trainer.train(wait=False)
# Container runs: pip install -r requirements.txt (using public PyPI, not CodeArtifact)
# Fails in VPC-isolated environments where PyPI is unreachable
v2 behaviour — the PyTorch estimator used the container's native entrypoint, which invoked sagemaker-training-toolkit. The toolkit checked CA_REPOSITORY_ARN, ran aws codeartifact login --tool pip, then installed requirements. This was added in sagemaker-training-toolkit v4.7.0 and tracked in deep-learning-containers#2509.
Suggested fix: The INSTALL_REQUIREMENTS template in sagemaker.train.templates should check for CA_REPOSITORY_ARN and configure pip accordingly before installing, matching the behaviour of sagemaker-training-toolkit.
- Ngôn ngữ chính
- Python
- Star
- 2.3k
- Fork
- 1.3k
- Merge trung bình
- 3 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 87
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của aws/sagemaker-python-sdk
-
Cannot use spark_event_logs_s3_uri in PySparkProcessor jobCó thể đã có người làm @rsareddy0329 đã nhận 8 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
aws/sagemaker-python-sdk#6253 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] V3 Hyperparameter Tuning Pipeline page labelled "Download Data" in navigation due to missing title cellCó thể đã có người làm @admivsn đã nhận 36 ngày trước. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 93/100
aws/sagemaker-python-sdk#6232 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] ModelTrainer with no input channels emits InputDataConfig: [], which CreatePipeline rejects (min=1) — v2 omitted the keyCó thể đã có người làm @sagemaker-bot đã nhận 8 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
aws/sagemaker-python-sdk#6156 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
sagemaker-train should depend on mlflow-skinny, following sagemaker-mlflow 0.5.0Có thể đã có người làm @mohamedzeidan2021 đã nhận 9 ngày trước. Đang mở
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 72/100
aws/sagemaker-python-sdk#6152 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
ModelTrainer generates sm_train.sh with CRLF line endings on Windows causing training job failureCó thể đã có người làm @MohammedAlkindi đã nhận 27 ngày trước. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
aws/sagemaker-python-sdk#5904 · 1 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của aws/sagemaker-python-sdk
Issue tương tự
-
area/install reliability
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 83/100
FluidNumerics/fluid-walk-blocker#191 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
TransformerLensOrg/TransformerLens#1868 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
climate-analytics-lab/jax-gcm#1057 ·
Maintainer thường phản hồi trong vòng 1 ngày