Issues with merge factor attributes when merge all = TRUE
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
Línea de trabajo
Reproduce the two merge.data.table() examples and compare their factor attributes, then inspect the related rbind/rbindlist handling around src/rbindlist.c line 350. Check how unmatched rows and factor columns are stacked, and add regression coverage showing that source_field_name is retained for both full outer joins. Done means the custom attribute survives the unmatched-row case without changing factor levels or class.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
merge.data.table() appears to drop custom attributes from factor columns when all = TRUE requires adding an unmatched row.
library('data.table')
d_x <- data.table(
id = 1:2,
value = structure(
c(1L,2L),
levels = c("No","Yes"),
class = c('ordered','factor'),
source_field_name = 'Q1'
)
)
d_y1 <- data.table(id = 1:2)
d_y2 <- data.table(id = 1:3)
d_mrg1 <- merge(d_x, d_y1, by = 'id', all = TRUE)
d_mrg2 <- merge(d_x, d_y2, by = 'id', all = TRUE)
attributes(d_mrg1$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
#
# $source_field_name
# [1] "Q1"
attributes(d_mrg2$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
The only difference is that d_y2 contains an unmatched id = 3. When that row is present, the custom source_field_name attribute disappears.
I would expect the custom attribute to be retained in both cases. This seems to be specific to factor columns; custom attributes on other column types appear to survive the same operation.
I encountered this because two otherwise very similar full outer joins produced different attribute results depending on whether an unmatched row happened to be present.
From what I can tell this is related to the rbind rbindlist stacking factor issue I had expected rbindlist()..., and here.
d_x <- data.table(
id = 1:2,
z = structure(
factor(c('No','Yes'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
)
)
attributes(d_x$z)
## Simple d_y with no z field. ##
d_y <- data.table(id = 3L)
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field but not attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE)))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field with attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
The funny thing is that I had written my own workaround version of rbindlist called rbl_safer that used my own preferences regarding how factors should be handled, but it is a lot harder to do for merge because of the suffix wrapping, etc. I could still write it, but I wanted to makes sure we knew there were secondary consequences. Handling rbindlist factor attributes is indeed "trickier than I initially thought", so I don't want to presume anything.
In the short term we could take off the factor restriction within rbindlist line L350, because the levels get establish at the end anyway.
> sessionInfo()
R version 4.6.1 (2026-06-24 ucrt)
Platform: x86_64-w64-mingw32/x64
Running under: Windows 11 x64 (build 26300)
Matrix products: default
LAPACK version 3.12.1
locale:
[1] LC_COLLATE=English_United States.utf8 LC_CTYPE=English_United States.utf8 LC_MONETARY=English_United States.utf8 LC_NUMERIC=C
[5] LC_TIME=English_United States.utf8
time zone: America/New_York
tzcode source: internal
attached base packages:
[1] stats graphics grDevices utils datasets methods base
other attached packages:
[1] data.table_1.18.99
loaded via a namespace (and not attached):
[1] compiler_4.6.1 tools_4.6.1
- Lenguaje dominante
- R
- Estrellas
- 3.9k
- Forks
- 1.1k
- Merge medio
- 15 h 51 min
- PR fusionados (30 d)
- 3
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Rdatatable/data.table#7887 ·
-
test() doesn't distinguish plain NA_real_, NaNPosiblemente ocupada @MichaelChirico la tomó hace 67 días. Abiertoconsistency tests
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Rdatatable/data.table#7853 · 3 comentarios ·
-
HAVE_LONG_DOUBLE is conditioned on but never setPosiblemente ocupada @venom1204 la tomó hace 511 días. Abiertointernals
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
Rdatatable/data.table#6938 · 1 comentario ·
-
encoding fread
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
Rdatatable/data.table#5179 · 8 comentarios ·
-
documentation programming
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Rdatatable/data.table#3199 · 3 comentarios ·
Todos los issues de Rdatatable/data.table
Issues similares
-
Installed vignettes show blank tables/plots: self_contained: no drops lt assets on R CMD buildAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
internal-code
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
inbo/erl-butterflies-2025#26 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 67/100