PyTorch on Azure ML: problem accessing GPU with ACPT environment image
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- azure, python, pytorch
- Ambito
- cloud, machine-learning
Direzione di ricerca
Inizia da 20_image_build_log.txt intorno alla riga 3408 e da training_script.py intorno alla riga 22, quindi esamina la documentazione collegata relativa ad Azure ML ACPT e ai container GPU. Riproduci l’esecuzione usando l’immagine, le dimensioni del cluster e l’ambiente conda indicati e verifica che PyTorch venga importato correttamente e possa accedere alla GPU senza l’errore relativo al simbolo CUDA.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
In Azure Machine Learning, I am trying to set up a PyTorch 2.0.1 run on a curated ACPT image:
mcr.microsoft.com/azureml/curated/acpt-pytorch-2.0-cuda11.7:latest
(based on this and this documentation)
I use the image above on a GPU cluster with size STANDARD_NC6and this conda environment yaml:
name: compass-environment-simple
channels:
- anaconda
- pytorch
- nvidia
- conda-forge
dependencies:
- python==3.10.9
- pip
- pytorch==2.0.1
- torchvision
- torchaudio
- pytorch-cuda=11.7
- cudatoolkit=11.7
- pip:
- pandas==2.0.1
- numpy==1.24.3
- azure-ai-ml==1.7.2
- azureml-mlflow==1.51.0
- mlflow==2.3.2
- tensorboard==2.13.0
- tensorboardX==2.6
- matplotlib==3.7.1
- matplotlib-inline==0.1.6
- mizani==0.9.1
- plotnine==0.12.1
- seaborn==0.12.2
- scikit-learn==1.2.2
- scipy==1.10.1
- numpy==1.24.3
- plotly==5.14.1
- statsmodels==0.14.0
- ipython==8.14.0
- tqdm==4.65.0
- umap-learn==0.5.3
- imbalanced-learn==0.10.1
- gensim==4.3.1
- nltk==3.8.1
- python-dotenv==1.0.0
- openpyxl==3.1.2
- xlwt==1.3.0
- xlrd==2.0.1
- XlsxWriter==3.1.0
But I keep getting a couple of error messages and I cannot access the GPUs when running the PyTorch Python scripts.
-
while building the image
in20_image_build_log.txtI get a warning that the NVIDIA driver can't be detected:
WARNING: The NVIDIA Driver was not detected. GPU functionality will not be available.
(see the attached 20_image_build_log.txt at line 3408) -
when executing the Python script during a run
I get the following error message:
/bin/bash: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/libtinfo.so.6: no version information available (required by /bin/bash)
['/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python310.zip', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/lib-dynload', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/mpmath-1.2.1-py3.10.egg', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b']
Traceback (most recent call last):
File "/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd/training_script.py", line 22, in <module>
import torch
File "/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/__init__.py", line 229, in <module>
from torch._C import * # noqa: F403
ImportError: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so: undefined symbol: cudaGraphInstantiateWithFlags, version libcudart.so.11.0
What am I doing wrong?
Thanks for your help.
- Lingua principale
- Jupyter Notebook
- Stelle
- 4.4k
- Fork
- 2.6k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Azure/MachineLearningNotebooks
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
Azure/MachineLearningNotebooks#1975 · 1 commento ·
-
duplicates Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 68/100
Azure/MachineLearningNotebooks#1960 ·
-
machine Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 10/100
Azure/MachineLearningNotebooks#1987 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
Azure/MachineLearningNotebooks#1985 · 1 commento ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 35/100
Azure/MachineLearningNotebooks#1981 · 1 reazione ·
Tutte le issue di Azure/MachineLearningNotebooks
Issue simili
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's Apertabug ecr
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
hashicorp/go-azure-helpers#286 ·
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
palladius/rails8-app-on-gcp#141 · 1 commento ·