Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[azureml python sdk v2] access files in URI_FOLDER output after job has finished?

未关闭
#1,891 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
停滞
技术栈
azure, jupyter-notebook, python

调研方向

先从 issue 中展示的 Azure ML v2 SDK job 和 NodeOutput API 开始,然后将其与 MlflowClient 访问路径进行比较。在 job ivory_octopus_yd6by49kxf 完成后,记录或演示一种以编程方式列出、检索和下载 URI_FOLDER 文件的方法,即视为完成。

由索引模型根据 Issue 内容生成。

描述

I have a training job that persists some files in an URI_FOLDER output.
How can I access those through the v2 SDK API after the job has finished?

1. job setup

The output is set up like this in the command:

job = command(
    # ...
    outputs=dict(
        outputs=Output(type=AssetTypes.URI_FOLDER, mode='rw_mount'),
    ),
    command="python training_script.py " + 
            "--outputs_dir ${{outputs.outputs}} " +
            # ...other arguments...
)

This seems to work fine, the corresponding folder is mounted correctly and accessible in the training script.

2. training script

In the training script, I persist a dataframe like this:

parser.add_argument("--outputs_dir", dest="outputs_dir", default=DEFAULT_MODEL_DIR)
# ...
some_dataframe.to_csv(os.path.join(args.outputs_dir, 'some_dataframe.csv'), index=True)

This works fine.

3. resulting dataset

After the job has finished, the outputs are available as a dataset.
This is what is shown in Azure ML Studio in the "Overview" tab for job ivory_octopus_yd6by49kxf:
image

The dataset is successfully stored in the workspaceblobstore datastore. I checked it in the Azure ML Studio and it looks fine.

4. accessing the persisted data

After the job has finished, I access the run using a MlflowClient()

MLFLOW_TRACKING_URI = ml_client.workspaces.get(name=ml_client.workspace_name).mlflow_tracking_uri
mlflow.set_tracking_uri(MLFLOW_TRACKING_URI)
mlflow_client = MlflowClient()

mlflow_run = mlflow_client.get_run("ivory_octopus_yd6by49kxf")

or

run = ml_client.jobs.get('ivory_octopus_yd6by49kxf')
# returns NodeOutput class

How can I programmatically list / get / download the outputs connected to the job?

Thanks!

主要语言
Jupyter Notebook
星标
4.4k
派生
2.6k
PR 合并指标
30 天内没有已合并 PR

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

Azure/MachineLearningNotebooks 的其他 Issue

查看 Azure/MachineLearningNotebooks 的全部 Issue

相似的 Issue

更多 Data Engineering Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。