drop flag on a column depends on the order of add_column and add_processor
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
Start with packages/data-designer-config/src/data_designer/config/config_builder.py around add_processor, _resolve_drop_column_names, and _remove_processor_by_name; then read engine/validation.py's resolve_processor_dropped_columns and engine/compiler.py's seed-column handling. Reproduce the three call orders in the issue and run the config and engine package tests. Done means all orders produce the same drop flags, re-adding a column preserves the intended behavior, and the glob-related duplicate drop behavior is addressed or explicitly scoped.
索引モデルが issue の本文から書いたものです。
説明
Priority Level
Medium (Annoying but has workaround)
Describe the bug
Whether a column ends up marked drop=True depends on the order the builder methods were called in, and on whether the column was added twice.
add_processor sets the flag as a side effect, and only on columns that are already in self._column_configs:
if processor_config.processor_type == ProcessorType.DROP_COLUMNS:
for col in self._resolve_drop_column_names(processor_config.column_names):
self._column_configs[col].drop = True
packages/data-designer-config/src/data_designer/config/config_builder.py:404, with _resolve_drop_column_names (:430) silently skipping names it does not know yet.
Two ways the flag goes missing:
add_processorbeforeadd_column. The column does not exist yet, so nothing is marked.add_columnagain for the same column after the processor. That is an upsert, so it replaces the config object and the flag goes back to its default.
validate_drop_columns_processor does not catch either, because by then the column does exist. The engine documents the consequence in _mark_processor_dropped_seed_columns (engine/compiler.py:58), which closes the same hole for seed columns; ordinary columns have nothing equivalent.
Steps/Code to reproduce bug
On main at 0b8c669. Same intent, three call orders:
from data_designer.config.config_builder import DataDesignerConfigBuilder
from data_designer.config.column_configs import SamplerColumnConfig
from data_designer.config.column_types import SamplerType
from data_designer.config.sampler_params import UUIDSamplerParams, CategorySamplerParams
from data_designer.config.processors import DropColumnsProcessorConfig
keep = lambda: SamplerColumnConfig(name="keep", sampler_type=SamplerType.UUID, params=UUIDSamplerParams())
tmp = lambda: SamplerColumnConfig(
name="tmp", sampler_type=SamplerType.CATEGORY, params=CategorySamplerParams(values=["a", "b"])
)
flags = lambda b: {c.name: c.drop for c in b.build().columns}
a = DataDesignerConfigBuilder()
a.add_column(keep()); a.add_column(tmp())
a.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b = DataDesignerConfigBuilder()
b.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
b.add_column(keep()); b.add_column(tmp())
c = DataDesignerConfigBuilder()
c.add_column(keep()); c.add_column(tmp())
c.add_processor(DropColumnsProcessorConfig(name="d", column_names=["tmp"]))
c.add_column(tmp())
print("columns -> processor :", flags(a))
print("processor -> columns :", flags(b))
print("re-added column :", flags(c))
columns -> processor : {'keep': False, 'tmp': True}
processor -> columns : {'keep': False, 'tmp': False}
re-added column : {'keep': False, 'tmp': False}
End to end, sampler columns only so no model is needed, with RunConfig(preserve_dropped_columns=False):
columns -> processor: OK -> columns ['keep']
processor -> columns: DataDesignerProfilingError: 🛑 Error profiling dataset: Column 'tmp' not found in dataset
The failure lands at the very end, after the whole dataset has been generated. With the default preserve_dropped_columns=True both orders report ['keep'], because the profiler reads the dataset back through load_dataset_with_dropped_columns(), so the divergence is invisible and only the config is wrong.
Expected behavior
The same calls in any order produce the same config, and re-adding a column does not clear a flag a processor set.
Agent Diagnostic / Prior Investigation
There is a second, smaller thing in the same area. dataset_builder.py:1069 compares config.column_names to columns_to_drop literally, with no glob expansion, so "tmp_*" never matches ["tmp_a", "tmp_b"] and a default_drop_columns_processor is added on top of the explicit one. Data comes out right, the noise does not:
drop flags: {'keep': False, 'tmp_a': True, 'tmp_b': True}
[INFO] 🙈 Dropping columns: ['tmp_*']
[INFO] 🙈 Dropping columns: ['tmp_a', 'tmp_b']
[WARNING] ⚠️ Cannot drop column: `tmp_a` not found in the dataset.
[WARNING] ⚠️ Cannot drop column: `tmp_b` not found in the dataset.
final columns: ['keep']
Searched before filing. gh issue list --state open --limit 80: nothing. gh pr list --state open --limit 50: the 6 open PRs do not touch config_builder.py or dataset_builder.py. Searches for drop processor order, drop_columns processor, add_processor (#10 and #332, closed), preserve_dropped_columns profiler, Column not found in dataset profiler, default_drop_columns_processor glob, dropped columns (#690, #798, #79, #177, #4, closed and unrelated), dataset profiler, builder order drop. The nearest is closed #332, "DropColumnsProcessorConfig: not idempotent on re-run", which is about re-adding the processor (fixed with _remove_processor_by_name) and about the validator rejecting *__reasoning_content. Re-adding the column, and the call order, are not in it.
Baseline, per package: data-designer-config 659 passed, data-designer-engine 2346 passed.
Additional context
What I would do, if it sounds right to you: drop the side effect entirely, which also removes the need for the manual drop=False rollback in _remove_processor_by_name (:412), and resolve the flags once in build(), or next to _mark_processor_dropped_seed_columns in the compiler, using the existing resolve_processor_dropped_columns in engine/validation.py:285, which already expands globs. The three orders above then agree, re-adding a column stops losing the flag, and the same call fixes the glob mismatch at dataset_builder.py:1069.
That moves the flag from "set when the processor happens to be added" to "derived from the processors", which is a small behaviour change for anyone reading drop off a half-built builder, so I would rather hear your preference before writing it.
- 主要言語
- Python
- スター
- 2.3k
- フォーク
- 219
- 平均マージ
- 1日 21時間
- マージ済み PR(30日)
- 38
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA-NeMo/DataDesigner のほかの issue
-
docs: required_columns description is incomplete for LLM and multimodal columns対応中かも @nightcityblade が今日担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
NVIDIA-NeMo/DataDesigner#1002 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
NVIDIA-NeMo/DataDesigner#995 ·
メンテナーはふだん 1 日以内に返信
-
enforce \from future import annotations` via ruff FA102 rule`対応中かも @chethanuk が 28 日前に担当しました。 オープンtask
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA-NeMo/DataDesigner#760 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 63/100
NVIDIA-NeMo/DataDesigner#1001 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 36/100
NVIDIA-NeMo/DataDesigner#994 ·
メンテナーはふだん 1 日以内に返信
NVIDIA-NeMo/DataDesigner の issue をすべて見る
似ている issue
-
documentation
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
ansys/pydpf-core#3547 ·
メンテナーはふだん 1 日以内に返信
-
core
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
vectorize-io/hindsight#5457 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recording対応中かも @ktz03 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
volcengine/OpenViking#5806 ·
メンテナーはふだん 1 日以内に返信
-
HTML: <template> content is extracted as document text対応中かも @ryanmeowy が今日担当しました。 オープンbug html
難易度 1/5 1時間未満 初心者へのやさしさ 82/100
docling-project/docling#4714 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
APIv2 event data accepts a non-string reply and a NaN upper_bound対応中かも @awss1i が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
freedomofpress/securedrop#7946 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信