Better handling for unrecognized categorical levels?
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
Direzione di ricerca
Non sono indicati file sorgente, test o un obiettivo concreto di implementazione. Inizia esaminando gli esempi di codifica dell’issue e i riferimenti collegati, quindi chiarisci il comportamento desiderato per i livelli categoriali non osservati e per ciascuno schema di codifica; il lavoro è completo quando, prima dell’implementazione, sono stati concordati un design e l’ambito corrispondente.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
There's a request here to add an option so that when an unrecognized categorical level is encountered, it should be encoded as all-zeros, which is apparently similar to what scikit-learn's DictVectorizer does.
Technically this is something patsy could do. But AFAICT this would lead to terribly incorrect behavior in any kind of linear-ish model, and AFAIK linear-ish models are what one-hot encoding are for, so I don't understand what's going on here or why people want this, and I like to understand things before implementing them :-).
Specifically, the kind of issue I'm thinking about is... say you have a logistic regression model you're using to predict whether an apartment is occupied, with a model like occupancy ~ C(city) + bedrooms + baths. In this model, patsy will use treatment coding, so returning all-zeros for unrecognized cities is the same as predicting that they act just like whichever city was assigned as the reference category (probably the one that's first alphabetically). OK, but that's not what DictVectorizer does -- it always uses a full-rank encoding, so it's more like patsy's occupancy ~ 0 + C(city) + bedrooms + baths. Now in this model, the beta for each city gives something like the (logistic-transformed) mean occupancy for each city, and using all-zeros for unrecognized cities is equivalent to assuming that their mean occupancy is exactly 0 on the logistic scale, which is a terrible guess. You really want it to do something like... return a vector of [1/n, 1/n, ..., 1/n], so that you're assuming unseen cities have similar occupancy to the average of the seen cities. Of course high-frequency cities and low-frequency cities are probably different too...
And other categorical encodings (e.g. polynomial coding) are even more of a mess.
So I'm not sure what to do here, if anything.
I feel like 99% of the time if you have an open category like this and don't want to use some principle solution like bayesian non-parametrics, then you instead want to do something like... keep the top 100 categories and bin everything else into "other", so then you actually have training data on the "other" category that's plausibly representative of what you'll see later (because in both cases it's relatively low frequency items). I guess this is also something patsy could potentially provide helpers for, though maybe it's more of a pandas thing.
CC: @ameuller
- Lingua principale
- Python
- Stelle
- 990
- Fork
- 106
- Merge medio
- 7g 34m
- PR unite (30g)
- 1
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di pydata/patsy
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 55/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
Tutte le issue di pydata/patsy
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
PedestrianDynamics/pyFDS-Evac#343 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
theskumar/python-dotenv#708 ·
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 2 giorni
-
Docs Timedelta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
pandas-dev/pandas#69919 ·
I maintainer di solito rispondono entro 1 giorno
-
API documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
zephyrproject-rtos/west#1009 · 2 commenti ·
I maintainer di solito rispondono entro 3 giorni