Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

HyperparameterTuner drops content_type when converting InputData to Channel

クローズ
#5,632 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
48/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
停滞
技術スタック
aws, python

調査の方向性

sagemaker/train/tuner.py の _create_hyperparameter_tuning_job から始め、特に 1362-1373 行目を確認して、InputData から Channel への変換を追跡します。提供されている HyperparameterTuner のケースを XGBoost で再現します。生成された Channel が content_type を保持し、トレーニングジョブが validate_data_file_path エラーで失敗しなくなった時点で作業は完了です。

索引モデルが issue の本文から書いたものです。

説明

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Describe the bug
When HyperparameterTuner.tune() receives InputData objects as inputs, it converts them to Channel objects internally but drops the content_type field during conversion. This causes built-in algorithms (e.g., XGBoost) to fail with validate_data_file_path errors because the container doesn't know the data format.

To reproduce

from sagemaker.train.configs import InputData
from sagemaker.train.tuner import HyperparameterTuner

train_input = InputData(
    channel_name="train",
    data_source="s3://my-bucket/train/train.csv",
    content_type="csv",  # <-- this gets dropped
)

tuner = HyperparameterTuner(
    model_trainer=model_trainer,
    objective_metric_name="validation:auc",
    hyperparameter_ranges=hyperparameter_ranges,
    objective_type="Maximize",
    max_jobs=12,
    max_parallel_jobs=3,
    strategy="Bayesian",
)

tuner.tune(inputs=[train_input])
# All training jobs fail with:
# AlgorithmError: validate_data_file_path(train_path, content_type)

Root Cause
In sagemaker/train/tuner.py, the _create_hyperparameter_tuning_job method converts InputData → Channel without passing content_type:


# tuner.py lines 1362-1373
```python
if isinstance(inp, InputData):
    input_data_config.append(Channel(
        channel_name=inp.channel_name,
        data_source=DataSource(
            s3_data_source=S3DataSource(
                s3_data_type="S3Prefix",
                s3_uri=inp.data_source,
                s3_data_distribution_type="FullyReplicated"
            )
        )
        # content_type is missing here!
    ))

Suggested Fix

if isinstance(inp, InputData):
    input_data_config.append(Channel(
        channel_name=inp.channel_name,
        content_type=inp.content_type,  # <-- add this
        data_source=DataSource(
            s3_data_source=S3DataSource(
                s3_data_type="S3Prefix",
                s3_uri=inp.data_source,
                s3_data_distribution_type="FullyReplicated"
            )
        )
    ))

Workaround
Pass Channel objects directly instead of InputData:

from sagemaker.core.shapes import Channel, DataSource, S3DataSource

train_input = Channel(
    channel_name="train",
    content_type="csv",
    data_source=DataSource(
        s3_data_source=S3DataSource(
            s3_data_type="S3Prefix",
            s3_uri="s3://my-bucket/train/train.csv",
            s3_data_distribution_type="FullyReplicated",
        )
    ),
)

tuner.tune(inputs=[train_input])  # works correctly

Environment
SageMaker Python SDK version: 3.0.1
Python version: 3.12
Built-in algorithm: XGBoost 1.7-1

主要言語
Python
スター
2.3k
フォーク
1.3k
平均マージ
3日 4時間
マージ済み PR(30日)
70

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

aws/sagemaker-python-sdk のほかの issue

aws/sagemaker-python-sdk の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。