Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Primitive for normalizing (feature scaling) input data

Aperta
#82 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
pandas, python

Direzione di ricerca

Inizia esaminando le convenzioni esistenti per le primitive Python e il modo in cui vengono gestiti i pandas DataFrames in questo repository. Risolvi se l'intervallo target è [-1,1] oppure quello [0,1] della formula fornita, insieme al comportamento previsto di trimming, clipping e inplace. Il lavoro è completato quando la primitiva di normalizzazione e i relativi casi limite sono coperti dai test.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

new primitives Pending Review

I want to create a primitive for normalization of data, or feature scaling so that the input is rescaled to [-1,1]

The formula for rescaling is
X_rescaled = (X - X.min) / (X.max - X.min)

The primitive input arguments are:

  • data (pandas dataframe)
  • column (string): the column to be rescaled
  • trim_percentage (float): percentage from the bottom and top to trim
  • inplace=True: if false, creates a new column called rescaled_input

Here are some potential issues with the implementation:

  • the implementation will run on historic data. If this primitive were to be used in an online system, we would have to either implement dynamic rescaling of dataset or automatically flag values larger than min and max as anomalies. Or we can just clip the input and assign large values as -1 or 1. The last suggestion makes the most sense, in my opinion.

  • distribution of the data. The data may have an outlier (very large value, e.g. 1234) and then the remaining values would be somewhere between -10 and 10. This would result in bad rescaling. One solution is to trim the lowest and highest 1% values, but what if those were outliers?

Lingua principale
Python
Stelle
70
Fork
37
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di MLBazaar/MLPrimitives

Tutte le issue di MLBazaar/MLPrimitives

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.