PyTorch on Azure ML: problem accessing GPU with ACPT environment image

未关闭
#1,917 12 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
停滞
技术栈
azure, python, pytorch

调研方向

先查看 20_image_build_log.txt 第 3408 行附近和 training_script.py 第 22 行附近,然后查阅链接的 Azure ML ACPT 和 GPU 容器文档。使用列出的镜像、集群大小和 conda 环境重现此次运行,并验证 PyTorch 能够成功导入且可以访问 GPU,而不会出现 CUDA 符号错误。

由索引模型根据 Issue 内容生成。

描述

In Azure Machine Learning, I am trying to set up a PyTorch 2.0.1 run on a curated ACPT image:
mcr.microsoft.com/azureml/curated/acpt-pytorch-2.0-cuda11.7:latest
(based on this and this documentation)

I use the image above on a GPU cluster with size STANDARD_NC6and this conda environment yaml:

name: compass-environment-simple
channels:
  - anaconda
  - pytorch
  - nvidia
  - conda-forge
dependencies:
  - python==3.10.9
  - pip
  - pytorch==2.0.1
  - torchvision
  - torchaudio  
  - pytorch-cuda=11.7
  - cudatoolkit=11.7
  - pip:
      - pandas==2.0.1
      - numpy==1.24.3
      - azure-ai-ml==1.7.2
      - azureml-mlflow==1.51.0
      - mlflow==2.3.2
      - tensorboard==2.13.0
      - tensorboardX==2.6
      - matplotlib==3.7.1
      - matplotlib-inline==0.1.6
      - mizani==0.9.1
      - plotnine==0.12.1
      - seaborn==0.12.2
      - scikit-learn==1.2.2
      - scipy==1.10.1
      - numpy==1.24.3
      - plotly==5.14.1
      - statsmodels==0.14.0
      - ipython==8.14.0
      - tqdm==4.65.0
      - umap-learn==0.5.3  
      - imbalanced-learn==0.10.1
      - gensim==4.3.1
      - nltk==3.8.1
      - python-dotenv==1.0.0
      - openpyxl==3.1.2
      - xlwt==1.3.0
      - xlrd==2.0.1
      - XlsxWriter==3.1.0

But I keep getting a couple of error messages and I cannot access the GPUs when running the PyTorch Python scripts.

  1. while building the image
    in 20_image_build_log.txt I get a warning that the NVIDIA driver can't be detected:
    WARNING: The NVIDIA Driver was not detected. GPU functionality will not be available.
    (see the attached 20_image_build_log.txt at line 3408)

  2. when executing the Python script during a run
    I get the following error message:

/bin/bash: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/libtinfo.so.6: no version information available (required by /bin/bash)
['/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python310.zip', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/lib-dynload', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/mpmath-1.2.1-py3.10.egg', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b']
Traceback (most recent call last):
  File "/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd/training_script.py", line 22, in <module>
    import torch
  File "/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/__init__.py", line 229, in <module>
    from torch._C import *  # noqa: F403
ImportError: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so: undefined symbol: cudaGraphInstantiateWithFlags, version libcudart.so.11.0

What am I doing wrong?
Thanks for your help.

主要语言
Jupyter Notebook
星标
4.4k
派生
2.6k
PR 合并指标
30 天内没有已合并 PR

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

Azure/MachineLearningNotebooks 的其他 Issue

查看 Azure/MachineLearningNotebooks 的全部 Issue

相似的 Issue

更多 Cloud Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。