Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Deferring Evaluation of Terms

Aperta
#149 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
python
Ambito
tooling

Direzione di ricerca

Non sono indicati file di implementazione né test. Inizia esaminando i punti di ingresso di Patsy per il parsing delle formule e la gestione dei termini, quindi determina se i termini assorbiti possono essere intercettati prima della costruzione della matrice densa; il lavoro sarà considerato completato quando saranno definiti una sintassi concordata e un design per il parsing di tali termini senza materializzarli.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

I have written a function called AbsorbingLS that can absorb a large number (millions) of categorical variables or categorical interactions. It is implemented using a Frisch-Waugh-Lovell step where the categoricals are handled using scipy sparse matrices. I would like to add a formula interface. Suppose I have a function A() that indicates that a variable should be absorbed, is there any place to intervene in the formula parsing for a formula that looks like y ~ 1 + x + A(cat) + A(cat*x)?

I can use another syntax. In an instrumental variable regression I use the syntax y ~ 1 + x1 + x2 + [x3 ~ z1 + z2] which is used to determine the configuration of the 2 required regressions. This works fine since it is easy to parse the [] and then it is a couple of standard calls. This approach doesn't obviously work here since I must avoid creating any arrays. I could use a similar structure here, so something like y ~ 1 + x + {cat +cat*x} ({} for simplicity in parsing) which I would then need to find a good way to turn cat +cat*x into usable terms (w/o populating dense arrays).

Any suggestions on how I could write a formula where I could intercept it using something like the pseudocode

1. Patsy  parses to terms
2. I remove and terms that are absorbed, which are too large to express as dense arrays
3. Patsy parses the non-absorbed terms, which have reasonable sizes as dense matrices

Any suggestions are appreciated.

Lingua principale
Python
Stelle
990
Fork
106
Merge medio
7g 34m
PR unite (30g)
1

Preparare l'ambiente

Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di pydata/patsy

Tutte le issue di pydata/patsy

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.