PyTorch on Azure ML: problem accessing GPU with ACPT environment image
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- azure, python, pytorch
- Domain
- cloud, machine-learning
Research direction
Start with 20_image_build_log.txt around line 3408 and training_script.py around line 22, then review the linked Azure ML ACPT and GPU container documentation. Reproduce the run using the listed image, cluster size, and conda environment, and verify that PyTorch imports successfully and can access the GPU without the CUDA symbol error.
Written by the indexing model from the issue text.
Description
In Azure Machine Learning, I am trying to set up a PyTorch 2.0.1 run on a curated ACPT image:
mcr.microsoft.com/azureml/curated/acpt-pytorch-2.0-cuda11.7:latest
(based on this and this documentation)
I use the image above on a GPU cluster with size STANDARD_NC6and this conda environment yaml:
name: compass-environment-simple
channels:
- anaconda
- pytorch
- nvidia
- conda-forge
dependencies:
- python==3.10.9
- pip
- pytorch==2.0.1
- torchvision
- torchaudio
- pytorch-cuda=11.7
- cudatoolkit=11.7
- pip:
- pandas==2.0.1
- numpy==1.24.3
- azure-ai-ml==1.7.2
- azureml-mlflow==1.51.0
- mlflow==2.3.2
- tensorboard==2.13.0
- tensorboardX==2.6
- matplotlib==3.7.1
- matplotlib-inline==0.1.6
- mizani==0.9.1
- plotnine==0.12.1
- seaborn==0.12.2
- scikit-learn==1.2.2
- scipy==1.10.1
- numpy==1.24.3
- plotly==5.14.1
- statsmodels==0.14.0
- ipython==8.14.0
- tqdm==4.65.0
- umap-learn==0.5.3
- imbalanced-learn==0.10.1
- gensim==4.3.1
- nltk==3.8.1
- python-dotenv==1.0.0
- openpyxl==3.1.2
- xlwt==1.3.0
- xlrd==2.0.1
- XlsxWriter==3.1.0
But I keep getting a couple of error messages and I cannot access the GPUs when running the PyTorch Python scripts.
-
while building the image
in20_image_build_log.txtI get a warning that the NVIDIA driver can't be detected:
WARNING: The NVIDIA Driver was not detected. GPU functionality will not be available.
(see the attached 20_image_build_log.txt at line 3408) -
when executing the Python script during a run
I get the following error message:
/bin/bash: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/libtinfo.so.6: no version information available (required by /bin/bash)
['/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python310.zip', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/lib-dynload', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/mpmath-1.2.1-py3.10.egg', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b']
Traceback (most recent call last):
File "/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd/training_script.py", line 22, in <module>
import torch
File "/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/__init__.py", line 229, in <module>
from torch._C import * # noqa: F403
ImportError: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so: undefined symbol: cudaGraphInstantiateWithFlags, version libcudart.so.11.0
What am I doing wrong?
Thanks for your help.
- Dominant language
- Jupyter Notebook
- Stars
- 4.4k
- Forks
- 2.6k
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Azure/MachineLearningNotebooks
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Azure/MachineLearningNotebooks#1975 · 1 comment ·
-
duplicates Open
Difficulty 1/5 Under an hour Newbie friendliness 68/100
Azure/MachineLearningNotebooks#1960 ·
-
machine Open
Difficulty 5/5 Over a week Newbie friendliness 10/100
Azure/MachineLearningNotebooks#1987 ·
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
Azure/MachineLearningNotebooks#1985 · 1 comment ·
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
Azure/MachineLearningNotebooks#1981 · 1 reaction ·
All issues in Azure/MachineLearningNotebooks
Similar issues
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's Openbug ecr
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
hashicorp/go-azure-helpers#286 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
palladius/rails8-app-on-gcp#141 · 1 comment ·