Issues with merge factor attributes when merge all = TRUE
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 52/100
Direzione di ricerca
Reproduce the two merge.data.table() examples and compare their factor attributes, then inspect the related rbind/rbindlist handling around src/rbindlist.c line 350. Check how unmatched rows and factor columns are stacked, and add regression coverage showing that source_field_name is retained for both full outer joins. Done means the custom attribute survives the unmatched-row case without changing factor levels or class.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
merge.data.table() appears to drop custom attributes from factor columns when all = TRUE requires adding an unmatched row.
library('data.table')
d_x <- data.table(
id = 1:2,
value = structure(
c(1L,2L),
levels = c("No","Yes"),
class = c('ordered','factor'),
source_field_name = 'Q1'
)
)
d_y1 <- data.table(id = 1:2)
d_y2 <- data.table(id = 1:3)
d_mrg1 <- merge(d_x, d_y1, by = 'id', all = TRUE)
d_mrg2 <- merge(d_x, d_y2, by = 'id', all = TRUE)
attributes(d_mrg1$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
#
# $source_field_name
# [1] "Q1"
attributes(d_mrg2$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
The only difference is that d_y2 contains an unmatched id = 3. When that row is present, the custom source_field_name attribute disappears.
I would expect the custom attribute to be retained in both cases. This seems to be specific to factor columns; custom attributes on other column types appear to survive the same operation.
I encountered this because two otherwise very similar full outer joins produced different attribute results depending on whether an unmatched row happened to be present.
From what I can tell this is related to the rbind rbindlist stacking factor issue I had expected rbindlist()..., and here.
d_x <- data.table(
id = 1:2,
z = structure(
factor(c('No','Yes'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
)
)
attributes(d_x$z)
## Simple d_y with no z field. ##
d_y <- data.table(id = 3L)
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field but not attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE)))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field with attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
The funny thing is that I had written my own workaround version of rbindlist called rbl_safer that used my own preferences regarding how factors should be handled, but it is a lot harder to do for merge because of the suffix wrapping, etc. I could still write it, but I wanted to makes sure we knew there were secondary consequences. Handling rbindlist factor attributes is indeed "trickier than I initially thought", so I don't want to presume anything.
In the short term we could take off the factor restriction within rbindlist line L350, because the levels get establish at the end anyway.
> sessionInfo()
R version 4.6.1 (2026-06-24 ucrt)
Platform: x86_64-w64-mingw32/x64
Running under: Windows 11 x64 (build 26300)
Matrix products: default
LAPACK version 3.12.1
locale:
[1] LC_COLLATE=English_United States.utf8 LC_CTYPE=English_United States.utf8 LC_MONETARY=English_United States.utf8 LC_NUMERIC=C
[5] LC_TIME=English_United States.utf8
time zone: America/New_York
tzcode source: internal
attached base packages:
[1] stats graphics grDevices utils datasets methods base
other attached packages:
[1] data.table_1.18.99
loaded via a namespace (and not attached):
[1] compiler_4.6.1 tools_4.6.1
- Lingua principale
- R
- Stelle
- 3.9k
- Fork
- 1.1k
- Merge medio
- 15h 51m
- PR unite (30g)
- 3
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Rdatatable/data.table#7887 ·
-
test() doesn't distinguish plain NA_real_, NaNForse già presa @MichaelChirico l’ha presa 67 giorni fa. Apertaconsistency tests
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Rdatatable/data.table#7853 · 3 commenti ·
-
HAVE_LONG_DOUBLE is conditioned on but never setForse già presa @venom1204 l’ha presa 511 giorni fa. Apertainternals
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
Rdatatable/data.table#6938 · 1 commento ·
-
encoding fread
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
Rdatatable/data.table#5179 · 8 commenti ·
-
documentation programming
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Rdatatable/data.table#3199 · 3 commenti ·
Tutte le issue di Rdatatable/data.table
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
inbo/erl-butterflies-2025#26 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 67/100
-
broken link to libgit2Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
RConsortium/submissions-pilot7-synthetic-data#95 ·
I maintainer di solito rispondono entro 1 giorno
-
betweenAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100