Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

drop flag on a column depends on the order of add_column and add_processor

オープン
#996 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
53/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python
領域
backend

調査の方向性

Start with packages/data-designer-config/src/data_designer/config/config_builder.py around add_processor, _resolve_drop_column_names, and _remove_processor_by_name; then read engine/validation.py's resolve_processor_dropped_columns and engine/compiler.py's seed-column handling. Reproduce the three call orders in the issue and run the config and engine package tests. Done means all orders produce the same drop flags, re-adding a column preserves the intended behavior, and the glob-related duplicate drop behavior is addressed or explicitly scoped.

索引モデルが issue の本文から書いたものです。

説明

Priority Level

Medium (Annoying but has workaround)

Describe the bug

Whether a column ends up marked drop=True depends on the order the builder methods were called in, and on whether the column was added twice.

add_processor sets the flag as a side effect, and only on columns that are already in self._column_configs:

if processor_config.processor_type == ProcessorType.DROP_COLUMNS:
    for col in self._resolve_drop_column_names(processor_config.column_names):
        self._column_configs[col].drop = True

packages/data-designer-config/src/data_designer/config/config_builder.py:404, with _resolve_drop_column_names (:430) silently skipping names it does not know yet.

Two ways the flag goes missing:

  1. add_processor before add_column. The column does not exist yet, so nothing is marked.
  2. add_column again for the same column after the processor. That is an upsert, so it replaces the config object and the flag goes back to its default.

validate_drop_columns_processor does not catch either, because by then the column does exist. The engine documents the consequence in _mark_processor_dropped_seed_columns (engine/compiler.py:58), which closes the same hole for seed columns; ordinary columns have nothing equivalent.

Steps/Code to reproduce bug

On main at 0b8c669. Same intent, three call orders:

from data_designer.config.config_builder import DataDesignerConfigBuilder
from data_designer.config.column_configs import SamplerColumnConfig
from data_designer.config.column_types import SamplerType
from data_designer.config.sampler_params import UUIDSamplerParams, CategorySamplerParams
from data_designer.config.processors import DropColumnsProcessorConfig

keep = lambda: SamplerColumnConfig(name="keep", sampler_type=SamplerType.UUID, params=UUIDSamplerParams())
tmp = lambda: SamplerColumnConfig(
    name="tmp", sampler_type=SamplerType.CATEGORY, params=CategorySamplerParams(values=["a", "b"])
)
flags = lambda b: {c.name: c.drop for c in b.build().columns}

a = DataDesignerConfigBuilder()
a.add_column(keep()); a.add_column(tmp())
a.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))

b = DataDesignerConfigBuilder()
b.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b.add_column(keep()); b.add_column(tmp())

c = DataDesignerConfigBuilder()
c.add_column(keep()); c.add_column(tmp())
c.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
c.add_column(tmp())

print("columns -> processor :", flags(a))
print("processor -> columns :", flags(b))
print("re-added column      :", flags(c))
columns -> processor : {'keep': False, 'tmp': True}
processor -> columns : {'keep': False, 'tmp': False}
re-added column      : {'keep': False, 'tmp': False}

End to end, sampler columns only so no model is needed, with RunConfig(preserve_dropped_columns=False):

columns -> processor: OK -> columns ['keep']
processor -> columns: DataDesignerProfilingError: 🛑 Error profiling dataset: Column 'tmp' not found in dataset

The failure lands at the very end, after the whole dataset has been generated. With the default preserve_dropped_columns=True both orders report ['keep'], because the profiler reads the dataset back through load_dataset_with_dropped_columns(), so the divergence is invisible and only the config is wrong.

Expected behavior

The same calls in any order produce the same config, and re-adding a column does not clear a flag a processor set.

Agent Diagnostic / Prior Investigation

There is a second, smaller thing in the same area. dataset_builder.py:1069 compares config.column_names to columns_to_drop literally, with no glob expansion, so "tmp_*" never matches ["tmp_a", "tmp_b"] and a default_drop_columns_processor is added on top of the explicit one. Data comes out right, the noise does not:

drop flags: {'keep': False, 'tmp_a': True, 'tmp_b': True}
[INFO] 🙈 Dropping columns: ['tmp_*']
[INFO] 🙈 Dropping columns: ['tmp_a', 'tmp_b']
[WARNING] ⚠️ Cannot drop column: `tmp_a` not found in the dataset.
[WARNING] ⚠️ Cannot drop column: `tmp_b` not found in the dataset.
final columns: ['keep']

Searched before filing. gh issue list --state open --limit 80: nothing. gh pr list --state open --limit 50: the 6 open PRs do not touch config_builder.py or dataset_builder.py. Searches for drop processor order, drop_columns processor, add_processor (#10 and #332, closed), preserve_dropped_columns profiler, Column not found in dataset profiler, default_drop_columns_processor glob, dropped columns (#690, #798, #79, #177, #4, closed and unrelated), dataset profiler, builder order drop. The nearest is closed #332, "DropColumnsProcessorConfig: not idempotent on re-run", which is about re-adding the processor (fixed with _remove_processor_by_name) and about the validator rejecting *__reasoning_content. Re-adding the column, and the call order, are not in it.

Baseline, per package: data-designer-config 659 passed, data-designer-engine 2346 passed.

Additional context

What I would do, if it sounds right to you: drop the side effect entirely, which also removes the need for the manual drop=False rollback in _remove_processor_by_name (:412), and resolve the flags once in build(), or next to _mark_processor_dropped_seed_columns in the compiler, using the existing resolve_processor_dropped_columns in engine/validation.py:285, which already expands globs. The three orders above then agree, re-adding a column stops losing the flag, and the same call fixes the glob mismatch at dataset_builder.py:1069.

That moves the flag from "set when the processor happens to be added" to "derived from the processors", which is a small behaviour change for anyone reading drop off a half-built builder, so I would rather hear your preference before writing it.

主要言語
Python
スター
2.3k
フォーク
219
平均マージ
1日 21時間
マージ済み PR(30日)
38

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA-NeMo/DataDesigner のほかの issue

NVIDIA-NeMo/DataDesigner の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。