Unable to run 3_evaluation.ipynb and 4_clusterability.ipynb correctly
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 25/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Ferma
- Stack tecnologico
- jupyter-notebook, python
- Ambito
- data, machine-learning
Direzione di ricerca
Start by comparing 2_train.ipynb, 3_evaluation.ipynb, and 4_clusterability.ipynb with the Code Ocean version, then inspect the trained_model value in pyproject.toml and the saving_folder paths. Re-run the training and evaluation notebooks while checking whether the expected 43 models, summary file, clustering folder, and .p files are produced. Done means both notebooks run with the documented inputs and produce the files their later cells load.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi, I'm kang. Thank you for sharing MMIDAS models.It is a very cool model! I'm trying to reproduce it lately. I had some problems running the code from github, so I ended up using the jupyter notebook downloaded from code ocean. But there are still some problems I can't solve.
The orignal code is copied from code ocean.
Problem 1:Unable to plot average consensus measure correctly
3_evaluation.ipynb:
In [5]:
#Loading trained models including before pruning and after pruning
cplMixVAE.variational = False
#summary_dict = summarize_inference(cplMixVAE, trained_models, trainloader)
To make the code work, I removed the # from the last line and added saving_folder after trainloader, and finally I changed the value of trained_model in pyproject.toml to match the real filename. So, the final edited code like this:
edited code:
#Loading trained models including before pruning and after pruning
cplMixVAE.variational = False
summary_dict = summarize_inference(cplMixVAE,trained_models,trainloader,saving_folder)
Than I run it and it works. I soon realized that only 4 models (1 before pruning + 3 after pruning) were recognized instead of 43 (1 before pruning + 42 after pruning). I checked the 2_train.ipynb file and found that:
2_train.ipynb
In [2]:
...............
max_prun_it = 3 # maximum number of pruning iterations
...............
So I modified max_prun_it to 42. So, the final edited code like this:
edited code:
max_prun_it = 42 # maximum number of pruning iterations
After retraining and running, all 43 models were detected, yielding the same results as jupyter notebook in code ocean.
Next I ran the following code:
3_evaluation.ipynb:
In [6]:
# Plotting average consensus measure to select the number of clusters according to the minimum consensus measure, here 0.95
summary_file = saving_folder + f'/summary_performance_K_{n_categories}_narm_{n_arm}.p'
with open(summary_file, 'rb') as f:
summary_dict = np.load(f, allow_pickle=True)
f.close()
model_order = K_selection(summary_dict, n_categories, n_arm, thr=0.95)
output:
Required minimum consensus is set too high, kindly consider specifying a lower value.
This's my biggest problem. It is inconsistent with the reference results and I'm not sure whether this is normal or not.
Problem 2: Unable to find .p files and folder called clustering
4_clusterability
ln[5]:
date = '20231121'
data_file_id = saving_folder + f"/clustering/Ttype_classification_K_{n_ttype}_nFeature_100_{date}.p"
sum_dict = pickle.load(open(data_file_id, "rb"))
acc_T_pc = sum_dict['acc_T_adj']
sc_T_pc = sum_dict['sc_T']
conf_T_pc = sum_dict['conf_mat']
.........................................
.........................................
Here “data_file_id=” specifies the path of the file, but immediately after running ”sum_dict = pickle.load(open(data_file_id, “rb”))“, an error occurred:
> FileNotFoundError: [Errno 2] No such file or directory:"here is the path of data_file_id"
When I checked my saving folders I realized that there was no folder called clustering, and there was no .p file that data_file_id was pointing to. I'm not sure if I'm missing any steps.
- Lingua principale
- Jupyter Notebook
- Stelle
- 6
- Fork
- 2
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
DOI-USGS/pywatershed#421 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
The-Strategy-Unit/nhp_output_reports#116 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
TUDelftGeodesy/DePSI#134 ·