Issues with merge factor attributes when merge all = TRUE
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 52/100
Direção de pesquisa
Reproduce the two merge.data.table() examples and compare their factor attributes, then inspect the related rbind/rbindlist handling around src/rbindlist.c line 350. Check how unmatched rows and factor columns are stacked, and add regression coverage showing that source_field_name is retained for both full outer joins. Done means the custom attribute survives the unmatched-row case without changing factor levels or class.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
merge.data.table() appears to drop custom attributes from factor columns when all = TRUE requires adding an unmatched row.
library('data.table')
d_x <- data.table(
id = 1:2,
value = structure(
c(1L,2L),
levels = c("No","Yes"),
class = c('ordered','factor'),
source_field_name = 'Q1'
)
)
d_y1 <- data.table(id = 1:2)
d_y2 <- data.table(id = 1:3)
d_mrg1 <- merge(d_x, d_y1, by = 'id', all = TRUE)
d_mrg2 <- merge(d_x, d_y2, by = 'id', all = TRUE)
attributes(d_mrg1$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
#
# $source_field_name
# [1] "Q1"
attributes(d_mrg2$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
The only difference is that d_y2 contains an unmatched id = 3. When that row is present, the custom source_field_name attribute disappears.
I would expect the custom attribute to be retained in both cases. This seems to be specific to factor columns; custom attributes on other column types appear to survive the same operation.
I encountered this because two otherwise very similar full outer joins produced different attribute results depending on whether an unmatched row happened to be present.
From what I can tell this is related to the rbind rbindlist stacking factor issue I had expected rbindlist()..., and here.
d_x <- data.table(
id = 1:2,
z = structure(
factor(c('No','Yes'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
)
)
attributes(d_x$z)
## Simple d_y with no z field. ##
d_y <- data.table(id = 3L)
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field but not attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE)))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field with attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
The funny thing is that I had written my own workaround version of rbindlist called rbl_safer that used my own preferences regarding how factors should be handled, but it is a lot harder to do for merge because of the suffix wrapping, etc. I could still write it, but I wanted to makes sure we knew there were secondary consequences. Handling rbindlist factor attributes is indeed "trickier than I initially thought", so I don't want to presume anything.
In the short term we could take off the factor restriction within rbindlist line L350, because the levels get establish at the end anyway.
> sessionInfo()
R version 4.6.1 (2026-06-24 ucrt)
Platform: x86_64-w64-mingw32/x64
Running under: Windows 11 x64 (build 26300)
Matrix products: default
LAPACK version 3.12.1
locale:
[1] LC_COLLATE=English_United States.utf8 LC_CTYPE=English_United States.utf8 LC_MONETARY=English_United States.utf8 LC_NUMERIC=C
[5] LC_TIME=English_United States.utf8
time zone: America/New_York
tzcode source: internal
attached base packages:
[1] stats graphics grDevices utils datasets methods base
other attached packages:
[1] data.table_1.18.99
loaded via a namespace (and not attached):
[1] compiler_4.6.1 tools_4.6.1
- Linguagem predominante
- R
- Estrelas
- 3.9k
- Forks
- 1.1k
- Merge médio
- 15h 51min
- PRs com merge (30d)
- 3
Preparar o ambiente
Inicia o contêiner de desenvolvimento do projeto no navegador, com a sua própria conta do GitHub.
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one)Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
Rdatatable/data.table#7887 ·
-
test() doesn't distinguish plain NA_real_, NaNTalvez já em andamento @MichaelChirico assumiu há 68 dias. Abertaconsistency tests
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Rdatatable/data.table#7853 · 3 comentários ·
-
HAVE_LONG_DOUBLE is conditioned on but never setTalvez já em andamento @venom1204 assumiu há 512 dias. Abertainternals
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
Rdatatable/data.table#6938 · 1 comentário ·
-
encoding fread
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
Rdatatable/data.table#5179 · 8 comentários ·
-
documentation programming
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Rdatatable/data.table#3199 · 3 comentários ·
Todas as issues de Rdatatable/data.table
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 76/100
datacarpentry/semester-biology#1272 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 75/100
DOI-USGS/dataRetrieval#934 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
-
pkgdown build failureAberta
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100