“EXPLAIN / profiling” support for [i, j, by] to diagnose performance bottlenecks
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Da chiarire
- Stato di attività
- Ferma
- Stack tecnologico
- r
- Ambito
- data, performance
Direzione di ricerca
The issue names no files, tests, or entry points. Start by reviewing how DT[i, j, by] handles filtering, grouping, joins, keys or indices, materialization, and copies. Define the diagnostics and acceptance criteria for the proposed explain or profiling mode.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
A common difficulty for users is understanding why a particular data.table query is slow.
When writing expressions like:
DT[ i , j , by]
there is currently no way to determine:
Whether time is spent in i filtering, j computation, or by grouping
Whether a key / index is actually being used
Whether grouping is triggering expensive materialization
Whether unexpected memory copies are occurring
Whether a join is using a fast path or falling back to a slower path
Most users rely on:
system.time(DT[i, j, by])
which measures only the total time and gives no insight into the internal bottleneck.
This makes optimization and debugging of complex workflows difficult, especially for users familiar with SQL-style tools like EXPLAIN.
Proposed Idea
Introduce an optional profiling / explain mode for data.table operations, conceptually similar to SQL’s EXPLAIN.
For example:
explain( DT[x > 5, .(m = mean(y)), by = z] )
or
options(datatable.explain = TRUE) DT[x > 5, .(m = mean(y)), by = z]
This could output structured diagnostics such as:
data.table EXPLAIN
Rows scanned: 5,000,000
Rows matched in i: 1,240,532
Groups formed by by: 2,134
Time spent:
i (filter): 120 ms
by (grouping): 340 ms
j (compute): 90 ms
Keys / indices used: YES (key: z)
Materialization: NO
Memory copies: 1 shallow copy
- Lingua principale
- R
- Stelle
- 3.9k
- Fork
- 1.1k
- Merge medio
- 15h 51m
- PR unite (30g)
- 3
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Rdatatable/data.table#7887 ·
-
test() doesn't distinguish plain NA_real_, NaNForse già presa @MichaelChirico l’ha presa 68 giorni fa. Apertaconsistency tests
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Rdatatable/data.table#7853 · 3 commenti ·
-
HAVE_LONG_DOUBLE is conditioned on but never setForse già presa @venom1204 l’ha presa 512 giorni fa. Apertainternals
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
Rdatatable/data.table#6938 · 1 commento ·
-
encoding fread
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
Rdatatable/data.table#5179 · 8 commenti ·
-
documentation programming
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Rdatatable/data.table#3199 · 3 commenti ·
Tutte le issue di Rdatatable/data.table
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
datacarpentry/semester-biology#1272 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 75/100
DOI-USGS/dataRetrieval#934 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
pkgdown build failureAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100