Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[v3] FrameworkProcessor and ModelTrainer: 4 regressions (including dropping CodeArtifact support) from v2 migration

Aperta
#5,765 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 2 giorni

Nessuno ha ancora preso questa issue.

  • #5790 di @aviruthen — chiusa senza merge

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
35/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
aws, python, pytorch

Direzione di ricerca

Inizia da sagemaker/core/resources.py per la gestione delle sessioni, sagemaker-core/src/sagemaker/core/processing.py per _package_code e _generate_framework_script e sagemaker.train.templates per INSTALL_REQUIREMENTS. Confronta il comportamento di CodeArtifact di v2 in PR #4145 con il comportamento di training-toolkit. Il lavoro è completo quando tutti e quattro i percorsi di regressione conservano la sessione fornita, rispettano code_location e installano i requisiti privati tramite CodeArtifact.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Additional context

The following four issues were found during a migration from sagemaker==2.257.1 to sagemaker==3.7.1, moving training jobs from sagemaker.pytorch.PyTorch to sagemaker.train.ModelTrainer and processing jobs from sagemaker.pytorch.processing.PyTorchProcessor to sagemaker.core.processing.FrameworkProcessor.


Describe the bug

Four regressions found when migrating from v2 to v3, specifically around FrameworkProcessor (processing jobs) and ModelTrainer (training jobs). All four worked correctly in v2.

System information

  • SageMaker Python SDK version: 3.7.1
  • Framework name: PyTorch
  • Framework version: 2.10
  • Python version: 3.13
  • CPU or GPU: Both
  • Custom Docker image: N

Bug 1: wait=True does not respect sagemaker session

Affects: ModelTrainer.train(wait=True) and FrameworkProcessor.run(wait=True)

ProcessingJob.refresh() and TrainingJob.refresh() use Base.get_sagemaker_client() — a global/default client — instead of the sagemaker_session passed to the processor/trainer. This fails with NoCredentialsError when using assumed-role sessions (via STS).

To reproduce:

import boto3
from sagemaker.core.helper.session_helper import Session
from sagemaker.core.processing import FrameworkProcessor
from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

# Assumed-role session
sts = boto3.client("sts")
assumed = sts.assume_role(RoleArn="arn:aws:iam::123456789:role/MyRole", RoleSessionName="test")
creds = assumed["Credentials"]
assumed_session = boto3.Session(
    aws_access_key_id=creds["AccessKeyId"],
    aws_secret_access_key=creds["SecretAccessKey"],
    aws_session_token=creds["SessionToken"],
    region_name="us-west-2",
)
sm_session = Session(boto_session=assumed_session)

# FrameworkProcessor — job created OK, wait fails
processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    sagemaker_session=sm_session,
)
processor.run(code="my_script.py", source_dir="src", wait=True)
# → NoCredentialsError

# ModelTrainer — same issue
trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(entry_script="train.py", source_dir="src"),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    sagemaker_session=sm_session,
)
trainer.train(wait=True)
# → NoCredentialsError

Root cause (sagemaker/core/resources.py):

# ProcessingJob.refresh() / TrainingJob.refresh()
client = Base.get_sagemaker_client()  # ← ignores the session
response = client.describe_processing_job(**operation_input_args)

v2 behaviour — ProcessingJob.wait() used the session directly:

def wait(self, logs=True):
    if logs:
        self.sagemaker_session.logs_for_processing_job(self.job_name, wait=True)
    else:
        self.sagemaker_session.wait_for_processing_job(self.job_name)

Bug 2: FrameworkProcessor.code_location is accepted but ignored

FrameworkProcessor.__init__ accepts code_location and stores it as self.code_location. The docstring states it controls where code is uploaded. However, _package_code ignores it and always uploads to self.sagemaker_session.default_bucket().

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
    code_location="s3://my-custom-bucket",  # ← ignored
)
processor.run(code="my_script.py", source_dir="src", wait=False)
# Code uploads to s3://sagemaker-us-west-2-123456789/... instead of s3://my-custom-bucket/...

Root cause (sagemaker-core/src/sagemaker/core/processing.py, _package_code):

s3_uri = s3.s3_path_join(
    "s3://",
    self.sagemaker_session.default_bucket(),  # ← always uses default bucket
    self.sagemaker_session.default_bucket_prefix or "",
    job_name, "source", "sourcedir.tar.gz",
)

self.code_location is never referenced.

v2 behaviour — FrameworkProcessor delegated to an estimator that honored code_location:

# v2 FrameworkProcessor._create_estimator
return self.estimator_cls(
    ...
    code_location=self.code_location,
    ...
)

Note: ModelTrainer does not offer code_location at all — it always uses session.default_bucket(). Suggested fix: either remove code_location from FrameworkProcessor to align with ModelTrainer, or update _package_code to use it when set.


Bugs 3 and 4 are both about losing the ability to install requirements.txt dependencies from a CodeArtifact repository.

  1. PyTorch training and inference containers supported this via the CA_REPOSITORY_ARN environment variable (see https://github.com/aws/deep-learning-containers/issues/2509 for details)
  2. https://github.com/aws/sagemaker-python-sdk/pull/4145 extended that support to processing jobs by exposing codeartifact_repo_arn on FrameworkProcessor.run()

For context, the new ray-based inference containers are also adding codeartifact support via CA_REPOSITORY_ARN environment variable as can be seen in https://github.com/aws/deep-learning-containers/blob/0fc07f317a4db68ff728274070fbe332dde1ca26/scripts/ray/sagemaker_serve.py#L68


Bug 3: CodeArtifact support missing from FrameworkProcessor

In v2, FrameworkProcessor.run() accepted codeartifact_repo_arn (added in PR #4145). This configured pip inside the container to authenticate with CodeArtifact before installing requirements.txt.

In v3, FrameworkProcessor.run() does not accept codeartifact_repo_arn, and _generate_framework_script has no CodeArtifact support. The generated runproc.sh runs pip install -r requirements.txt without authentication, failing for packages hosted on private CodeArtifact repositories.

To reproduce:

processor = FrameworkProcessor(
    image_uri="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    command=["python3"],
    role="arn:aws:iam::123456789:role/MyRole",
    instance_count=1,
    instance_type="ml.m5.xlarge",
)

# v2 supported: processor.run(..., codeartifact_repo_arn="arn:aws:codeartifact:us-west-2:123:repository/domain/repo")
# v3 does not — parameter doesn't exist

processor.run(code="my_script.py", source_dir="src", wait=False)
# Container fails: pip install -r requirements.txt → package not found (private CodeArtifact repo)

v2 behaviour — _generate_framework_script injected CodeArtifact login into runproc.sh:

if [[ -f 'requirements.txt' ]]; then
    if ! hash aws 2>/dev/null; then
        echo "AWS CLI is not installed. Skipping CodeArtifact login."
    else
        aws codeartifact login --tool pip --domain {domain} --domain-owner {owner} --repository {repository} --region {region}
    fi
    pip install -r requirements.txt
fi

Suggested fix: Port codeartifact_repo_arn and _get_codeartifact_command from v2's PR #4145 into v3's FrameworkProcessor.


Bug 4: ModelTrainer bypasses sagemaker-training-toolkit, losing CodeArtifact support for requirements.txt

Affects: ModelTrainer.train() with SourceCode(requirements="requirements.txt")

ModelTrainer overrides the container's ENTRYPOINT with its own sm_train.sh driver script. This bypasses the sagemaker-training-toolkit installed in the container, which handled requirements.txt installation with CodeArtifact support (via the CA_REPOSITORY_ARN environment variable, added in sagemaker-training-toolkit#187).

The v3 sm_train.sh template for requirements installation is a bare pip install:

# from sagemaker.train.templates.INSTALL_REQUIREMENTS
echo "Installing requirements"
$SM_PIP_CMD install -r {requirements_file}

It does not check for CA_REPOSITORY_ARN or configure pip to use CodeArtifact. The env var is passed to the container but nothing reads it.

To reproduce:

from sagemaker.train import ModelTrainer
from sagemaker.core.training.configs import Compute, SourceCode

trainer = ModelTrainer(
    training_image="763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-training:2.10-cpu-py313",
    role="arn:aws:iam::123456789:role/MyRole",
    source_code=SourceCode(
        entry_script="train.py",
        source_dir="src",
        requirements="requirements.txt",  # ← installed without CodeArtifact
    ),
    compute=Compute(instance_type="ml.m5.xlarge", instance_count=1),
    environment={"CA_REPOSITORY_ARN": "arn:aws:codeartifact:us-west-2:123:repository/domain/repo"},
)
trainer.train(wait=False)
# Container runs: pip install -r requirements.txt (using public PyPI, not CodeArtifact)
# Fails in VPC-isolated environments where PyPI is unreachable

v2 behaviour — the PyTorch estimator used the container's native entrypoint, which invoked sagemaker-training-toolkit. The toolkit checked CA_REPOSITORY_ARN, ran aws codeartifact login --tool pip, then installed requirements. This was added in sagemaker-training-toolkit v4.7.0 and tracked in deep-learning-containers#2509.

Suggested fix: The INSTALL_REQUIREMENTS template in sagemaker.train.templates should check for CA_REPOSITORY_ARN and configure pip accordingly before installing, matching the behaviour of sagemaker-training-toolkit.

Lingua principale
Python
Stelle
2.3k
Fork
1.3k
Merge medio
3g 10h
PR unite (30g)
77

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di aws/sagemaker-python-sdk

Tutte le issue di aws/sagemaker-python-sdk

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.