HashSlap-Summer-of-Code/ml-core

Standardize Dataset Handling Across All ML Modules

開放

#11 建立於 2025年6月18日

 (0 則留言) (0 個反應) (0 位負責人)Jupyter Notebook (11 個分叉)auto 404
Intermediateenhancementgood first issuehacktoberfesthssoc

倉庫指標

星標
 (3 顆星)
PR 合併指標
 (30 天內沒有已合併 PR)

描述

Description:
Right now, different subfolders load and process datasets in inconsistent ways. Create a unified, reusable Python module to handle dataset loading and basic preprocessing (e.g., scaling, splitting). This ensures maintainability and reduces repeated code across ML scripts.

Expected Tasks:

  • Create a Python utility (e.g., data_utils.py) with functions like:
    • load_csv(path)
    • train_test_split(X, y, test_size=0.2)
    • standardize(X)
  • Save this file in a utils/ folder.
  • Refactor at least two existing implementations (e.g., KNN and Perceptron) to use this utility.
  • Update documentation in the root README.md and/or affected folders.

Stretch Goal:

  • Add optional support for downloading public datasets (e.g., from UCI or sklearn).

貢獻者指南