HashSlap-Summer-of-Code/ml-core

Standardize Dataset Handling Across All ML Modules

开放

#11 创建于 2025年6月18日

 (0 条评论) (0 个反应) (0 位负责人)Jupyter Notebook (11 个派生)auto 404
Intermediateenhancementgood first issuehacktoberfesthssoc

仓库指标

星标
 (3 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

Description:
Right now, different subfolders load and process datasets in inconsistent ways. Create a unified, reusable Python module to handle dataset loading and basic preprocessing (e.g., scaling, splitting). This ensures maintainability and reduces repeated code across ML scripts.

Expected Tasks:

  • Create a Python utility (e.g., data_utils.py) with functions like:
    • load_csv(path)
    • train_test_split(X, y, test_size=0.2)
    • standardize(X)
  • Save this file in a utils/ folder.
  • Refactor at least two existing implementations (e.g., KNN and Perceptron) to use this utility.
  • Update documentation in the root README.md and/or affected folders.

Stretch Goal:

  • Add optional support for downloading public datasets (e.g., from UCI or sklearn).

贡献者指南