Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Numeric sampler rounding can violate constraints and invalidate conditional sampling

オープン
#1,001 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
63/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
data

調査の方向性

Read packages/data-designer-engine/src/data_designer/engine/sampling_gen/generator.py, especially _run_rejection_sampling() and generate(), and inspect the generator and constraint tests. Reproduce the rounding cases from the issue, then add regression coverage for constraints and conditional dependencies using final rounded values. Done means returned values satisfy configured constraints and dependent conditions reflect the values in the returned dataset, including the unsatisfiable-after-rounding case.

索引モデルが issue の本文から書いたものです。

説明

Priority Level

Medium (Annoying but has workaround)

Describe the bug

Setting decimal_places on a numeric sampler can silently produce final values that violate an explicit ScalarInequalityConstraint. Conditional samplers can also select a branch inconsistent with the final value of their source column.

Both behaviors arise because decimal rounding happens after rejection sampling and after dependent columns have been generated. Generation succeeds, but the returned data violates its configured rules.

Steps/Code to reproduce bug

Reproduced on an unmodified checkout of main at 0b8c669c717ed4806a96f8dd3629868900045dc5, using Python 3.10.12 and dependencies installed with uv sync --all-packages --frozen. No model calls are needed.

from data_designer.config import ScalarInequalityConstraint
from data_designer.engine.sampling_gen.generator import DatasetGenerator
from data_designer.engine.sampling_gen.schema import DataSchema

schema = DataSchema(
    columns=[{
        "name": "value",
        "sampler_type": "uniform",
        "params": {"low": 0, "high": 1, "decimal_places": 0},
    }],
    constraints=[ScalarInequalityConstraint(
        target_column="value", operator="gt", rhs=0.25,
    )],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(256)

print((frame["value"] <= 0.25).sum())  # 78
assert (frame["value"] > 0.25).all()  # fails

Actual: 78 of 256 final values are 0, violating value > 0.25, without an error. Changing only decimal_places to None yields zero violations.

A second reproduction demonstrates inconsistent conditional labels:

schema = DataSchema(
    columns=[
        {
            "name": "value",
            "sampler_type": "uniform",
            "params": {"low": 0.41, "high": 0.49, "decimal_places": 0},
        },
        {
            "name": "label",
            "sampler_type": "category",
            "params": {"values": ["low"]},
            "conditional_params": {
                "value > 0.4": {"values": ["high"]},
            },
        },
    ],
    constraints=[],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(16)

expected = frame["value"].apply(lambda value: "high" if value > 0.4 else "low")
print(frame.head())
print((frame["label"] != expected).sum())  # 16

All 16 records contain value=0.0 and label="high", although the final source values make that condition false. Changing decimal_places to None yields no contradictory labels. This conditional case was also reproduced through DataDesigner.create() and load_dataset().

Expected behavior

Every returned numeric value satisfies the configured constraint. The first configuration is feasible: rounded values of 1 satisfy value > 0.25; values that round to 0 should be rejected and resampled.

Conditional samplers should evaluate the source values that will appear in the resulting dataset, so the second example should produce label="low".

Agent Diagnostic / Prior Investigation

Investigation and reproductions were performed with a coding agent.

In generator.py:

  1. _run_rejection_sampling() runs preproc and then checks constraints.
  2. generate() completes sampling for all columns, including dependent conditional columns.
  3. Only afterward does it call _round_if_needed().

Consequently, constraints and dependent conditions see pre-rounding values that differ from the returned dataset.

Existing generator tests assert constraints and conditional relationships on final returned values, while precision tests cover decimal_places independently. The sampling documentation does not describe an exception allowing rounded output to violate constraints.

Validation against the unmodified checkout:

  • Existing test_generator.py and test_constraints.py: 35 passed.
  • Four additional behavioral checks: both unrounded controls pass; the rounded constraint and conditional cases fail (2 passed, 2 failed).
  • Reproduction imports were verified to resolve to the unmodified checkout.

Duplicate check on 2026-10-08 covered all 50 open issues and 5 open PRs, the 100 most recently updated closed PRs, historical issue/PR searches for precision/rounding/constraints/conditional sampling, and repository documentation and plans. No matching report or active PR was found. Nearby #484/#512 concern datetime formatting; #414 concerns constraint discriminators; #773 concerns workflow-level repeat-until stages.

Additional context

A proposed fix direction is to apply configured numeric rounding before checking constraints and before downstream samplers evaluate conditions, consistently on each rejection-sampling attempt. Datetime/string representation postprocessing should remain separate.

Regression coverage should include scalar and column inequalities, numeric conditional dependencies, and constraints that become unsatisfiable after rounding.

I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.

Checklist
  • I reproduced this issue or provided a minimal example
  • I searched the docs/issues myself, or had my agent do so
  • If I used an agent, I included its diagnostics above
主要言語
Python
スター
2.3k
フォーク
219
平均マージ
1日 21時間
マージ済み PR(30日)
38

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA-NeMo/DataDesigner のほかの issue

NVIDIA-NeMo/DataDesigner の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。