Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Numeric sampler rounding can violate constraints and invalidate conditional sampling

Open
#1,001 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
63/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
data

Research direction

Read packages/data-designer-engine/src/data_designer/engine/sampling_gen/generator.py, especially _run_rejection_sampling() and generate(), and inspect the generator and constraint tests. Reproduce the rounding cases from the issue, then add regression coverage for constraints and conditional dependencies using final rounded values. Done means returned values satisfy configured constraints and dependent conditions reflect the values in the returned dataset, including the unsatisfiable-after-rounding case.

Written by the indexing model from the issue text.

Description

Priority Level

Medium (Annoying but has workaround)

Describe the bug

Setting decimal_places on a numeric sampler can silently produce final values that violate an explicit ScalarInequalityConstraint. Conditional samplers can also select a branch inconsistent with the final value of their source column.

Both behaviors arise because decimal rounding happens after rejection sampling and after dependent columns have been generated. Generation succeeds, but the returned data violates its configured rules.

Steps/Code to reproduce bug

Reproduced on an unmodified checkout of main at 0b8c669c717ed4806a96f8dd3629868900045dc5, using Python 3.10.12 and dependencies installed with uv sync --all-packages --frozen. No model calls are needed.

from data_designer.config import ScalarInequalityConstraint
from data_designer.engine.sampling_gen.generator import DatasetGenerator
from data_designer.engine.sampling_gen.schema import DataSchema

schema = DataSchema(
    columns=[{
        "name": "value",
        "sampler_type": "uniform",
        "params": {"low": 0, "high": 1, "decimal_places": 0},
    }],
    constraints=[ScalarInequalityConstraint(
        target_column="value", operator="gt", rhs=0.25,
    )],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(256)

print((frame["value"] <= 0.25).sum())  # 78
assert (frame["value"] > 0.25).all()  # fails

Actual: 78 of 256 final values are 0, violating value > 0.25, without an error. Changing only decimal_places to None yields zero violations.

A second reproduction demonstrates inconsistent conditional labels:

schema = DataSchema(
    columns=[
        {
            "name": "value",
            "sampler_type": "uniform",
            "params": {"low": 0.41, "high": 0.49, "decimal_places": 0},
        },
        {
            "name": "label",
            "sampler_type": "category",
            "params": {"values": ["low"]},
            "conditional_params": {
                "value > 0.4": {"values": ["high"]},
            },
        },
    ],
    constraints=[],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(16)

expected = frame["value"].apply(lambda value: "high" if value > 0.4 else "low")
print(frame.head())
print((frame["label"] != expected).sum())  # 16

All 16 records contain value=0.0 and label="high", although the final source values make that condition false. Changing decimal_places to None yields no contradictory labels. This conditional case was also reproduced through DataDesigner.create() and load_dataset().

Expected behavior

Every returned numeric value satisfies the configured constraint. The first configuration is feasible: rounded values of 1 satisfy value > 0.25; values that round to 0 should be rejected and resampled.

Conditional samplers should evaluate the source values that will appear in the resulting dataset, so the second example should produce label="low".

Agent Diagnostic / Prior Investigation

Investigation and reproductions were performed with a coding agent.

In generator.py:

  1. _run_rejection_sampling() runs preproc and then checks constraints.
  2. generate() completes sampling for all columns, including dependent conditional columns.
  3. Only afterward does it call _round_if_needed().

Consequently, constraints and dependent conditions see pre-rounding values that differ from the returned dataset.

Existing generator tests assert constraints and conditional relationships on final returned values, while precision tests cover decimal_places independently. The sampling documentation does not describe an exception allowing rounded output to violate constraints.

Validation against the unmodified checkout:

  • Existing test_generator.py and test_constraints.py: 35 passed.
  • Four additional behavioral checks: both unrounded controls pass; the rounded constraint and conditional cases fail (2 passed, 2 failed).
  • Reproduction imports were verified to resolve to the unmodified checkout.

Duplicate check on 2026-10-08 covered all 50 open issues and 5 open PRs, the 100 most recently updated closed PRs, historical issue/PR searches for precision/rounding/constraints/conditional sampling, and repository documentation and plans. No matching report or active PR was found. Nearby #484/#512 concern datetime formatting; #414 concerns constraint discriminators; #773 concerns workflow-level repeat-until stages.

Additional context

A proposed fix direction is to apply configured numeric rounding before checking constraints and before downstream samplers evaluate conditions, consistently on each rejection-sampling attempt. Datetime/string representation postprocessing should remain separate.

Regression coverage should include scalar and column inequalities, numeric conditional dependencies, and constraints that become unsatisfiable after rounding.

I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.

Checklist
  • I reproduced this issue or provided a minimal example
  • I searched the docs/issues myself, or had my agent do so
  • If I used an agent, I included its diagnostics above
Dominant language
Python
Stars
2.3k
Forks
219
Avg merge
1d 16h
Merged PRs (30d)
42

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA-NeMo/DataDesigner

All issues in NVIDIA-NeMo/DataDesigner

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.