drop flag on a column depends on the order of add_column and add_processor
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 53/100
Línea de trabajo
Start with packages/data-designer-config/src/data_designer/config/config_builder.py around add_processor, _resolve_drop_column_names, and _remove_processor_by_name; then read engine/validation.py's resolve_processor_dropped_columns and engine/compiler.py's seed-column handling. Reproduce the three call orders in the issue and run the config and engine package tests. Done means all orders produce the same drop flags, re-adding a column preserves the intended behavior, and the glob-related duplicate drop behavior is addressed or explicitly scoped.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Priority Level
Medium (Annoying but has workaround)
Describe the bug
Whether a column ends up marked drop=True depends on the order the builder methods were called in, and on whether the column was added twice.
add_processor sets the flag as a side effect, and only on columns that are already in self._column_configs:
if processor_config.processor_type == ProcessorType.DROP_COLUMNS:
for col in self._resolve_drop_column_names(processor_config.column_names):
self._column_configs[col].drop = True
packages/data-designer-config/src/data_designer/config/config_builder.py:404, with _resolve_drop_column_names (:430) silently skipping names it does not know yet.
Two ways the flag goes missing:
add_processorbeforeadd_column. The column does not exist yet, so nothing is marked.add_columnagain for the same column after the processor. That is an upsert, so it replaces the config object and the flag goes back to its default.
validate_drop_columns_processor does not catch either, because by then the column does exist. The engine documents the consequence in _mark_processor_dropped_seed_columns (engine/compiler.py:58), which closes the same hole for seed columns; ordinary columns have nothing equivalent.
Steps/Code to reproduce bug
On main at 0b8c669. Same intent, three call orders:
from data_designer.config.config_builder import DataDesignerConfigBuilder
from data_designer.config.column_configs import SamplerColumnConfig
from data_designer.config.column_types import SamplerType
from data_designer.config.sampler_params import UUIDSamplerParams, CategorySamplerParams
from data_designer.config.processors import DropColumnsProcessorConfig
keep = lambda: SamplerColumnConfig(name="keep", sampler_type=SamplerType.UUID, params=UUIDSamplerParams())
tmp = lambda: SamplerColumnConfig(
name="tmp", sampler_type=SamplerType.CATEGORY, params=CategorySamplerParams(values=["a", "b"])
)
flags = lambda b: {c.name: c.drop for c in b.build().columns}
a = DataDesignerConfigBuilder()
a.add_column(keep()); a.add_column(tmp())
a.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b = DataDesignerConfigBuilder()
b.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b.add_column(keep()); b.add_column(tmp())
c = DataDesignerConfigBuilder()
c.add_column(keep()); c.add_column(tmp())
c.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
c.add_column(tmp())
print("columns -> processor :", flags(a))
print("processor -> columns :", flags(b))
print("re-added column :", flags(c))
columns -> processor : {'keep': False, 'tmp': True}
processor -> columns : {'keep': False, 'tmp': False}
re-added column : {'keep': False, 'tmp': False}
End to end, sampler columns only so no model is needed, with RunConfig(preserve_dropped_columns=False):
columns -> processor: OK -> columns ['keep']
processor -> columns: DataDesignerProfilingError: 🛑 Error profiling dataset: Column 'tmp' not found in dataset
The failure lands at the very end, after the whole dataset has been generated. With the default preserve_dropped_columns=True both orders report ['keep'], because the profiler reads the dataset back through load_dataset_with_dropped_columns(), so the divergence is invisible and only the config is wrong.
Expected behavior
The same calls in any order produce the same config, and re-adding a column does not clear a flag a processor set.
Agent Diagnostic / Prior Investigation
There is a second, smaller thing in the same area. dataset_builder.py:1069 compares config.column_names to columns_to_drop literally, with no glob expansion, so "tmp_*" never matches ["tmp_a", "tmp_b"] and a default_drop_columns_processor is added on top of the explicit one. Data comes out right, the noise does not:
drop flags: {'keep': False, 'tmp_a': True, 'tmp_b': True}
[INFO] 🙈 Dropping columns: ['tmp_*']
[INFO] 🙈 Dropping columns: ['tmp_a', 'tmp_b']
[WARNING] ⚠️ Cannot drop column: `tmp_a` not found in the dataset.
[WARNING] ⚠️ Cannot drop column: `tmp_b` not found in the dataset.
final columns: ['keep']
Searched before filing. gh issue list --state open --limit 80: nothing. gh pr list --state open --limit 50: the 6 open PRs do not touch config_builder.py or dataset_builder.py. Searches for drop processor order, drop_columns processor, add_processor (#10 and #332, closed), preserve_dropped_columns profiler, Column not found in dataset profiler, default_drop_columns_processor glob, dropped columns (#690, #798, #79, #177, #4, closed and unrelated), dataset profiler, builder order drop. The nearest is closed #332, "DropColumnsProcessorConfig: not idempotent on re-run", which is about re-adding the processor (fixed with _remove_processor_by_name) and about the validator rejecting *__reasoning_content. Re-adding the column, and the call order, are not in it.
Baseline, per package: data-designer-config 659 passed, data-designer-engine 2346 passed.
Additional context
What I would do, if it sounds right to you: drop the side effect entirely, which also removes the need for the manual drop=False rollback in _remove_processor_by_name (:412), and resolve the flags once in build(), or next to _mark_processor_dropped_seed_columns in the compiler, using the existing resolve_processor_dropped_columns in engine/validation.py:285, which already expands globs. The three orders above then agree, re-adding a column stops losing the flag, and the same call fixes the glob mismatch at dataset_builder.py:1069.
That moves the flag from "set when the processor happens to be added" to "derived from the processors", which is a small behaviour change for anyone reading drop off a half-built builder, so I would rather hear your preference before writing it.
- Lenguaje dominante
- Python
- Estrellas
- 2.3k
- Forks
- 219
- Merge medio
- 1 d 21 h
- PR fusionados (30 d)
- 38
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA-NeMo/DataDesigner
-
docs: required_columns description is incomplete for LLM and multimodal columnsPosiblemente ocupada @nightcityblade la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
NVIDIA-NeMo/DataDesigner#1002 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
NVIDIA-NeMo/DataDesigner#995 ·
Los mantenedores suelen responder en 1 día
-
enforce \from future import annotations` via ruff FA102 rule`Posiblemente ocupada @chethanuk la tomó hace 28 días. Abiertotask
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 63/100
NVIDIA-NeMo/DataDesigner#1001 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 36/100
NVIDIA-NeMo/DataDesigner#994 ·
Los mantenedores suelen responder en 1 día
Todos los issues de NVIDIA-NeMo/DataDesigner
Issues similares
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Abiertodocumentation priority: low size: S
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Los mantenedores suelen responder en 1 día
-
--csv-bom was never wired up: PR #850 added an unused helper parameter, so #846 is not fixedAbiertobug help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Los mantenedores suelen responder en 1 día
-
Broken link in index.rstAbiertodocumentation
Dificultad 1/5 Menos de una hora Aptitud para principiantes 65/100
ansys/pydpf-core#3547 ·
Los mantenedores suelen responder en 1 día
-
core
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
vectorize-io/hindsight#5457 ·
Los mantenedores suelen responder en 1 día
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recordingPosiblemente ocupada @ktz03 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
volcengine/OpenViking#5806 ·
Los mantenedores suelen responder en 1 día