Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

drop flag on a column depends on the order of add_column and add_processor

Abierto
#996 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
53/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
python
Área
backend

Línea de trabajo

Start with packages/data-designer-config/src/data_designer/config/config_builder.py around add_processor, _resolve_drop_column_names, and _remove_processor_by_name; then read engine/validation.py's resolve_processor_dropped_columns and engine/compiler.py's seed-column handling. Reproduce the three call orders in the issue and run the config and engine package tests. Done means all orders produce the same drop flags, re-adding a column preserves the intended behavior, and the glob-related duplicate drop behavior is addressed or explicitly scoped.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Priority Level

Medium (Annoying but has workaround)

Describe the bug

Whether a column ends up marked drop=True depends on the order the builder methods were called in, and on whether the column was added twice.

add_processor sets the flag as a side effect, and only on columns that are already in self._column_configs:

if processor_config.processor_type == ProcessorType.DROP_COLUMNS:
    for col in self._resolve_drop_column_names(processor_config.column_names):
        self._column_configs[col].drop = True

packages/data-designer-config/src/data_designer/config/config_builder.py:404, with _resolve_drop_column_names (:430) silently skipping names it does not know yet.

Two ways the flag goes missing:

  1. add_processor before add_column. The column does not exist yet, so nothing is marked.
  2. add_column again for the same column after the processor. That is an upsert, so it replaces the config object and the flag goes back to its default.

validate_drop_columns_processor does not catch either, because by then the column does exist. The engine documents the consequence in _mark_processor_dropped_seed_columns (engine/compiler.py:58), which closes the same hole for seed columns; ordinary columns have nothing equivalent.

Steps/Code to reproduce bug

On main at 0b8c669. Same intent, three call orders:

from data_designer.config.config_builder import DataDesignerConfigBuilder
from data_designer.config.column_configs import SamplerColumnConfig
from data_designer.config.column_types import SamplerType
from data_designer.config.sampler_params import UUIDSamplerParams, CategorySamplerParams
from data_designer.config.processors import DropColumnsProcessorConfig

keep = lambda: SamplerColumnConfig(name="keep", sampler_type=SamplerType.UUID, params=UUIDSamplerParams())
tmp = lambda: SamplerColumnConfig(
    name="tmp", sampler_type=SamplerType.CATEGORY, params=CategorySamplerParams(values=["a", "b"])
)
flags = lambda b: {c.name: c.drop for c in b.build().columns}

a = DataDesignerConfigBuilder()
a.add_column(keep()); a.add_column(tmp())
a.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))

b = DataDesignerConfigBuilder()
b.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b.add_column(keep()); b.add_column(tmp())

c = DataDesignerConfigBuilder()
c.add_column(keep()); c.add_column(tmp())
c.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
c.add_column(tmp())

print("columns -> processor :", flags(a))
print("processor -> columns :", flags(b))
print("re-added column      :", flags(c))
columns -> processor : {'keep': False, 'tmp': True}
processor -> columns : {'keep': False, 'tmp': False}
re-added column      : {'keep': False, 'tmp': False}

End to end, sampler columns only so no model is needed, with RunConfig(preserve_dropped_columns=False):

columns -> processor: OK -> columns ['keep']
processor -> columns: DataDesignerProfilingError: 🛑 Error profiling dataset: Column 'tmp' not found in dataset

The failure lands at the very end, after the whole dataset has been generated. With the default preserve_dropped_columns=True both orders report ['keep'], because the profiler reads the dataset back through load_dataset_with_dropped_columns(), so the divergence is invisible and only the config is wrong.

Expected behavior

The same calls in any order produce the same config, and re-adding a column does not clear a flag a processor set.

Agent Diagnostic / Prior Investigation

There is a second, smaller thing in the same area. dataset_builder.py:1069 compares config.column_names to columns_to_drop literally, with no glob expansion, so "tmp_*" never matches ["tmp_a", "tmp_b"] and a default_drop_columns_processor is added on top of the explicit one. Data comes out right, the noise does not:

drop flags: {'keep': False, 'tmp_a': True, 'tmp_b': True}
[INFO] 🙈 Dropping columns: ['tmp_*']
[INFO] 🙈 Dropping columns: ['tmp_a', 'tmp_b']
[WARNING] ⚠️ Cannot drop column: `tmp_a` not found in the dataset.
[WARNING] ⚠️ Cannot drop column: `tmp_b` not found in the dataset.
final columns: ['keep']

Searched before filing. gh issue list --state open --limit 80: nothing. gh pr list --state open --limit 50: the 6 open PRs do not touch config_builder.py or dataset_builder.py. Searches for drop processor order, drop_columns processor, add_processor (#10 and #332, closed), preserve_dropped_columns profiler, Column not found in dataset profiler, default_drop_columns_processor glob, dropped columns (#690, #798, #79, #177, #4, closed and unrelated), dataset profiler, builder order drop. The nearest is closed #332, "DropColumnsProcessorConfig: not idempotent on re-run", which is about re-adding the processor (fixed with _remove_processor_by_name) and about the validator rejecting *__reasoning_content. Re-adding the column, and the call order, are not in it.

Baseline, per package: data-designer-config 659 passed, data-designer-engine 2346 passed.

Additional context

What I would do, if it sounds right to you: drop the side effect entirely, which also removes the need for the manual drop=False rollback in _remove_processor_by_name (:412), and resolve the flags once in build(), or next to _mark_processor_dropped_seed_columns in the compiler, using the existing resolve_processor_dropped_columns in engine/validation.py:285, which already expands globs. The three orders above then agree, re-adding a column stops losing the flag, and the same call fixes the glob mismatch at dataset_builder.py:1069.

That moves the flag from "set when the processor happens to be added" to "derived from the processors", which is a small behaviour change for anyone reading drop off a half-built builder, so I would rather hear your preference before writing it.

Lenguaje dominante
Python
Estrellas
2.3k
Forks
219
Merge medio
1 d 21 h
PR fusionados (30 d)
38

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de NVIDIA-NeMo/DataDesigner

Todos los issues de NVIDIA-NeMo/DataDesigner

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.