Issues with merge factor attributes when merge all = TRUE
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 52/100
Rechercherichtung
Reproduce the two merge.data.table() examples and compare their factor attributes, then inspect the related rbind/rbindlist handling around src/rbindlist.c line 350. Check how unmatched rows and factor columns are stacked, and add regression coverage showing that source_field_name is retained for both full outer joins. Done means the custom attribute survives the unmatched-row case without changing factor levels or class.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
merge.data.table() appears to drop custom attributes from factor columns when all = TRUE requires adding an unmatched row.
library('data.table')
d_x <- data.table(
id = 1:2,
value = structure(
c(1L,2L),
levels = c("No","Yes"),
class = c('ordered','factor'),
source_field_name = 'Q1'
)
)
d_y1 <- data.table(id = 1:2)
d_y2 <- data.table(id = 1:3)
d_mrg1 <- merge(d_x, d_y1, by = 'id', all = TRUE)
d_mrg2 <- merge(d_x, d_y2, by = 'id', all = TRUE)
attributes(d_mrg1$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
#
# $source_field_name
# [1] "Q1"
attributes(d_mrg2$value)
# $levels
# [1] "No" "Yes"
#
# $class
# [1] "ordered" "factor"
The only difference is that d_y2 contains an unmatched id = 3. When that row is present, the custom source_field_name attribute disappears.
I would expect the custom attribute to be retained in both cases. This seems to be specific to factor columns; custom attributes on other column types appear to survive the same operation.
I encountered this because two otherwise very similar full outer joins produced different attribute results depending on whether an unmatched row happened to be present.
From what I can tell this is related to the rbind rbindlist stacking factor issue I had expected rbindlist()..., and here.
d_x <- data.table(
id = 1:2,
z = structure(
factor(c('No','Yes'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
)
)
attributes(d_x$z)
## Simple d_y with no z field. ##
d_y <- data.table(id = 3L)
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field but not attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE)))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
## d_y with Factor z field with attribute. ##
d_y <- data.table(id = 3L, z = structure(
factor(c('No'), levels = c('No','Yes'), ordered = TRUE),
source_field_name = 'Q1'
))
d_z <- rbind(d_x, d_y, fill = TRUE)
attributes(d_z$z)
The funny thing is that I had written my own workaround version of rbindlist called rbl_safer that used my own preferences regarding how factors should be handled, but it is a lot harder to do for merge because of the suffix wrapping, etc. I could still write it, but I wanted to makes sure we knew there were secondary consequences. Handling rbindlist factor attributes is indeed "trickier than I initially thought", so I don't want to presume anything.
In the short term we could take off the factor restriction within rbindlist line L350, because the levels get establish at the end anyway.
> sessionInfo()
R version 4.6.1 (2026-06-24 ucrt)
Platform: x86_64-w64-mingw32/x64
Running under: Windows 11 x64 (build 26300)
Matrix products: default
LAPACK version 3.12.1
locale:
[1] LC_COLLATE=English_United States.utf8 LC_CTYPE=English_United States.utf8 LC_MONETARY=English_United States.utf8 LC_NUMERIC=C
[5] LC_TIME=English_United States.utf8
time zone: America/New_York
tzcode source: internal
attached base packages:
[1] stats graphics grDevices utils datasets methods base
other attached packages:
[1] data.table_1.18.99
loaded via a namespace (and not attached):
[1] compiler_4.6.1 tools_4.6.1
- Vorherrschende Sprache
- R
- Sterne
- 3.9k
- Forks
- 1.1k
- Ø Merge
- 15 Std. 51 Min.
- Gemergte PRs (30 T.)
- 3
Entwicklungsumgebung
Startet den Dev-Container des Projekts im Browser, mit Ihrem eigenen GitHub-Konto.
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one)Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Rdatatable/data.table#7887 ·
-
test() doesn't distinguish plain NA_real_, NaNEvtl. vergeben @MichaelChirico hat das vor 69 Tagen übernommen. Offenconsistency tests
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
Rdatatable/data.table#7853 · 3 Kommentare ·
-
HAVE_LONG_DOUBLE is conditioned on but never setEvtl. vergeben @venom1204 hat das vor 513 Tagen übernommen. Offeninternals
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
Rdatatable/data.table#6938 · 1 Kommentar ·
-
encoding fread
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
Rdatatable/data.table#5179 · 8 Kommentare ·
-
documentation programming
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
Rdatatable/data.table#3199 · 3 Kommentare ·
Alle Issues in Rdatatable/data.table
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
-
bug code cleaning
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
furrer-lab/abn#272 ·
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 92/100
philchalmers/SimDesign#106 ·
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 64/100
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 75/100
hturner/pkg-dev-ctv#57 ·