Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Pipeline parameters used with DataPath and DataPathComputeBinding to specify side inputs of Parallel pipeline

未關閉
#1,801 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
35/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
停滯
技術堆疊
azure, python

研究方向

從 PipelineParameter API 頁面和連結的 AzureML-Docset 原始檔開始,然後使用 azureml-core==1.40.0.post2 和 azureml-pipeline==1.40.0 重現該範例。將文件中 DataPath/DataPathComputeBinding 的用法與回報的 ParallelRunStep 錯誤進行比較;當文件或支援版本指南準確反映該行為時,即視為完成。

由索引模型根據 Issue 內容生成。

描述

[Enter feedback here]
I'm following this example to create a PipelineParameters for my Parallel pipeline

from azureml.core.datastore import Datastore
from azureml.data.datapath import DataPath, DataPathComputeBinding
from azureml.pipeline.steps import PythonScriptStep
from azureml.pipeline.core import PipelineParameter

datastore = Datastore(workspace=workspace, name="workspaceblobstore")
datapath = DataPath(datastore=datastore, path_on_datastore='input_data')
data_path_pipeline_param = (PipelineParameter(name="input_data", default_value=datapath),
                           DataPathComputeBinding(mode='mount'))

train_step = PythonScriptStep(script_name="train.py",
                             arguments=["--input", data_path_pipeline_param],
                             inputs=[data_path_pipeline_param],
                             compute_target=compute_target,
                             source_directory=project_folder)

This is my code to create the pipeline with the parameters

path = DataPath(datastore=default_store, path_on_datastore='path')
input_param= (PipelineParameter(name="param_name", default_value=path), DataPathComputeBinding(mode='mount'))

parallel_run_config = ParallelRunConfig(
    source_directory=script_dir,
    entry_script='script.py',  # the user script to run against each input
    partition_keys=['key'],
    error_threshold=50,
    output_action='append_row',
    environment=environment,
    compute_target=compute_target, 
    node_count=2,
    run_invocation_timeout=1200
)

parallel_run_step = ParallelRunStep(
    name='test-batch-inference',
    inputs=[partition_input],
    side_inputs=[input1, input2, input_param],
    output=output_dir,
    parallel_run_config=parallel_run_config,
    arguments=['--input_param', input_param],
    allow_reuse=False
)

And it raised this error:

Exception: Step input must be of any type: (<class 'azureml.data.dataset_consumption_config.DatasetConsumptionConfig'>, <class 'azureml.pipeline.core.pipeline_output_dataset.PipelineOutputFileDataset'>, <class 'azureml.pipeline.core.pipeline_output_dataset.PipelineOutputTabularDataset'>, <class 'azureml.data.output_dataset_config.OutputFileDatasetConfig'>, <class 'azureml.data.output_dataset_config.OutputTabularDatasetConfig'>, <class 'azureml.data.output_dataset_config.LinkFileOutputDatasetConfig'>, <class 'azureml.data.output_dataset_config.LinkTabularOutputDatasetConfig'>), found <class 'tuple'>

I'm using azureml-core==1.40.0.post2, azureml-pipeline==1.40.0
It's seems like the sample code is not supported with these version? Before trying this datapath as pipeline parameter, I tried int type input and its just work fine


Document Details

⚠ Do not edit this section. It is required for docs.microsoft.com ➟ GitHub issue linking.

主要語言
Jupyter Notebook
星號
4.4k
分支
2.6k
PR 合併指標
30 天內沒有已合併 PR

環境準備

這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

Azure/MachineLearningNotebooks 的其他 Issue

查看 Azure/MachineLearningNotebooks 的全部 Issue

相似的 Issue

更多 Documentation Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。