Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Unable to run 3_evaluation.ipynb and 4_clusterability.ipynb correctly

Aperta
#1 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
25/100
Tipo di issue
Bug
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
jupyter-notebook, python

Direzione di ricerca

Start by comparing 2_train.ipynb, 3_evaluation.ipynb, and 4_clusterability.ipynb with the Code Ocean version, then inspect the trained_model value in pyproject.toml and the saving_folder paths. Re-run the training and evaluation notebooks while checking whether the expected 43 models, summary file, clustering folder, and .p files are produced. Done means both notebooks run with the documented inputs and produce the files their later cells load.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Hi, I'm kang. Thank you for sharing MMIDAS models.It is a very cool model! I'm trying to reproduce it lately. I had some problems running the code from github, so I ended up using the jupyter notebook downloaded from code ocean. But there are still some problems I can't solve.

The orignal code is copied from code ocean.

Problem 1:Unable to plot average consensus measure correctly
3_evaluation.ipynb: 

In [5]:
#Loading trained models including before pruning and after pruning
cplMixVAE.variational = False
#summary_dict = summarize_inference(cplMixVAE, trained_models, trainloader)

To make the code work, I removed the # from the last line and added saving_folder after trainloader, and finally I changed the value of trained_model in pyproject.toml to match the real filename. So, the final edited code like this:

edited code:

#Loading trained models including before pruning and after pruning
cplMixVAE.variational = False
summary_dict = summarize_inference(cplMixVAE,trained_models,trainloader,saving_folder)

Than I run it and it works. I soon realized that only 4 models (1 before pruning + 3 after pruning) were recognized instead of 43 (1 before pruning + 42 after pruning). I checked the 2_train.ipynb file and found that:

2_train.ipynb
 
In [2]:
...............
max_prun_it = 3 # maximum number of pruning iterations
...............

So I modified max_prun_it to 42. So, the final edited code like this:

edited code:

max_prun_it = 42 # maximum number of pruning iterations

After retraining and running, all 43 models were detected, yielding the same results as jupyter notebook in code ocean.

Next I ran the following code:

3_evaluation.ipynb: 

In [6]:

# Plotting average consensus measure to select the number of clusters according to the minimum consensus measure, here 0.95
summary_file = saving_folder + f'/summary_performance_K_{n_categories}_narm_{n_arm}.p'
with open(summary_file, 'rb') as f:
    summary_dict = np.load(f, allow_pickle=True)
f.close()
model_order = K_selection(summary_dict, n_categories, n_arm, thr=0.95)

output:
Required minimum consensus is set too high, kindly consider specifying a lower value.
6bf088eb-20f3-4656-ba5b-62de31e72e63

This's my biggest problem. It is inconsistent with the reference results and I'm not sure whether this is normal or not.

Problem 2: Unable to find .p files and folder called clustering
4_clusterability

ln[5]:

date = '20231121' 
data_file_id = saving_folder + f"/clustering/Ttype_classification_K_{n_ttype}_nFeature_100_{date}.p" 
sum_dict = pickle.load(open(data_file_id, "rb")) 
acc_T_pc = sum_dict['acc_T_adj']
sc_T_pc = sum_dict['sc_T']
conf_T_pc = sum_dict['conf_mat']
.........................................
.........................................

Here “data_file_id=” specifies the path of the file, but immediately after running ”sum_dict = pickle.load(open(data_file_id, “rb”))“, an error occurred:

> FileNotFoundError: [Errno 2] No such file or directory:"here is the path of data_file_id"

When I checked my saving folders I realized that there was no folder called clustering, and there was no .p file that data_file_id was pointing to. I'm not sure if I'm missing any steps.

Lingua principale
Jupyter Notebook
Stelle
6
Fork
2
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Issue simili

Altre issue su Data Engineering

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.