Add an option to control the maximum marginal degree in AIM workload construction
I maintainer di solito rispondono entro 2 giorni
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 65/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- ai-infra-agents, data, machine-learning
Direzione di ricerca
Look at the AIM workload construction code, likely in a file like aim.py or workload.py. The current fixed degree of 3 needs to be made a configurable parameter. Understand how max_marginal_size is used to filter marginals. The DPMM implementation linked shows the pattern. Add a degree parameter, integrate it with the existing filtering, and ensure it works with the two-column case from issue #196. Test by constructing workloads with different degree limits.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
AIM currently exposes max_marginal_size, which limits the size of marginal queries considered when constructing the workload. However, the maximum degree of the automatically constructed default workload is currently fixed at 3 and is not exposed as a configurable hyperparameter.
This is also related to the two-column case addressed in #196, where the number of available columns in the dataset necessarily results in a smaller degree than the current default workload construction.
In the DPMM implementation of AIM, these are exposed as two separate controls: degree (default 2) controls the maximum number of columns in a marginal, while max_cells constrains the size of the resulting marginal. max_cells therefore serves a similar purpose to max_marginal_size here.
It would be useful to expose the same distinction in this implementation. degree could determine the maximum number of columns considered together, while max_marginal_size could continue to filter these marginals based on the size of their joint domain. For example, a degree-3 marginal over three binary columns has only 8 cells, whereas a degree-2 marginal over two high-cardinality columns may be much larger.
The choice of degree and max_marginal_size, and the interaction between them, is an interesting hyperparameter trade-off.
If this sounds useful, I’d be happy to implement it and submit a PR.
- Lingua principale
- Python
- Stelle
- 32
- Fork
- 13
- Merge medio
- 1g 19h
- PR unite (30g)
- 20
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di google/dpsynth
-
import dpsynth fails because mbi.Dataset is registered as a JAX dataclass twiceForse già presa @hanzalaareeb l’ha presa 9 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 2 giorni
-
Clarify installation requirements in quickstart.ipynbForse già presa @hanzalaareeb l’ha presa 8 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
google/dpsynth#194 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
`IndependentConfig` synthesis raises "Cliques must be unique."Forse di nuovo libera Una pull request per questa issue è stata chiusa senza essere unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 2 giorni
-
Windows install of pylock.toml fails on the pipeline extra due to missing Windows wheel for python-dpForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 72/100
I maintainer di solito rispondono entro 2 giorni
-
Lazy-load TensorFlow for TFRecord-specific pipeline pathsForse già presa @hanzalaareeb l’ha presa 48 giorni fa. Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 68/100
google/dpsynth#151 · 2 commenti ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di google/dpsynth
Issue simili
-
Harmony OPeNDAP SubSetter (HOSS) Geographic LARC_CLOUD PREFIRE_SAT2_AUX-SAT R01 production
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
nasa/harmony-autotester#245 ·
-
feature
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
-
L: github:actions L: php:composer
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
dependabot/dependabot-core#16493 ·
I maintainer di solito rispondono entro 1 giorno
-
2.3 EDA: `np.log` example adds +1 to every value, so it can't produce the `-inf` output shownAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
DataTalksClub/machine-learning-zoomcamp#730 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 4 giorni