Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Training my own dataset

Open
#20 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
30/100
Issue type
Documentation
Clarity
Needs clarification
Activity status
Stale

Research direction

Start by reviewing the data\cath directory and the issue's distinction between PDB files and the expected Dataset format. Document the required dataset structure, how compatibility is determined, and the steps for training on a custom dataset. Done means a newcomer can prepare compatible data and follow the training process without needing an unanswered explanation.

Written by the indexing model from the issue text.

Description

Hi, thank you for sharing your very interesting folding model!
I want to train a model using my own dataset. Do I just need to put the dataset in the data\cath directory? Also, I would appreciate it if you could explain how to create a compatible Dataset. It seems like PDB files are not compatible.

Dominant language
Jupyter Notebook
Stars
568
Forks
74
Avg merge
20h 2m
Merged PRs (30d)
1

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/foldingdiff

All issues in microsoft/foldingdiff

Similar issues

More Data Engineering issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.