Numeric sampler rounding can violate constraints and invalidate conditional sampling
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 63/100
Research direction
Read packages/data-designer-engine/src/data_designer/engine/sampling_gen/generator.py, especially _run_rejection_sampling() and generate(), and inspect the generator and constraint tests. Reproduce the rounding cases from the issue, then add regression coverage for constraints and conditional dependencies using final rounded values. Done means returned values satisfy configured constraints and dependent conditions reflect the values in the returned dataset, including the unsatisfiable-after-rounding case.
Written by the indexing model from the issue text.
Description
Priority Level
Medium (Annoying but has workaround)
Describe the bug
Setting decimal_places on a numeric sampler can silently produce final values that violate an explicit ScalarInequalityConstraint. Conditional samplers can also select a branch inconsistent with the final value of their source column.
Both behaviors arise because decimal rounding happens after rejection sampling and after dependent columns have been generated. Generation succeeds, but the returned data violates its configured rules.
Steps/Code to reproduce bug
Reproduced on an unmodified checkout of main at 0b8c669c717ed4806a96f8dd3629868900045dc5, using Python 3.10.12 and dependencies installed with uv sync --all-packages --frozen. No model calls are needed.
from data_designer.config import ScalarInequalityConstraint
from data_designer.engine.sampling_gen.generator import DatasetGenerator
from data_designer.engine.sampling_gen.schema import DataSchema
schema = DataSchema(
columns=[{
"name": "value",
"sampler_type": "uniform",
"params": {"low": 0, "high": 1, "decimal_places": 0},
}],
constraints=[ScalarInequalityConstraint(
target_column="value", operator="gt", rhs=0.25,
)],
)
frame = DatasetGenerator(
sampler_columns=None, schema=schema, random_state=42,
).generate(256)
print((frame["value"] <= 0.25).sum()) # 78
assert (frame["value"] > 0.25).all() # fails
Actual: 78 of 256 final values are 0, violating value > 0.25, without an error. Changing only decimal_places to None yields zero violations.
A second reproduction demonstrates inconsistent conditional labels:
schema = DataSchema(
columns=[
{
"name": "value",
"sampler_type": "uniform",
"params": {"low": 0.41, "high": 0.49, "decimal_places": 0},
},
{
"name": "label",
"sampler_type": "category",
"params": {"values": ["low"]},
"conditional_params": {
"value > 0.4": {"values": ["high"]},
},
},
],
constraints=[],
)
frame = DatasetGenerator(
sampler_columns=None, schema=schema, random_state=42,
).generate(16)
expected = frame["value"].apply(lambda value: "high" if value > 0.4 else "low")
print(frame.head())
print((frame["label"] != expected).sum()) # 16
All 16 records contain value=0.0 and label="high", although the final source values make that condition false. Changing decimal_places to None yields no contradictory labels. This conditional case was also reproduced through DataDesigner.create() and load_dataset().
Expected behavior
Every returned numeric value satisfies the configured constraint. The first configuration is feasible: rounded values of 1 satisfy value > 0.25; values that round to 0 should be rejected and resampled.
Conditional samplers should evaluate the source values that will appear in the resulting dataset, so the second example should produce label="low".
Agent Diagnostic / Prior Investigation
Investigation and reproductions were performed with a coding agent.
In generator.py:
_run_rejection_sampling()runspreprocand then checks constraints.generate()completes sampling for all columns, including dependent conditional columns.- Only afterward does it call
_round_if_needed().
Consequently, constraints and dependent conditions see pre-rounding values that differ from the returned dataset.
Existing generator tests assert constraints and conditional relationships on final returned values, while precision tests cover decimal_places independently. The sampling documentation does not describe an exception allowing rounded output to violate constraints.
Validation against the unmodified checkout:
- Existing
test_generator.pyandtest_constraints.py: 35 passed. - Four additional behavioral checks: both unrounded controls pass; the rounded constraint and conditional cases fail (2 passed, 2 failed).
- Reproduction imports were verified to resolve to the unmodified checkout.
Duplicate check on 2026-10-08 covered all 50 open issues and 5 open PRs, the 100 most recently updated closed PRs, historical issue/PR searches for precision/rounding/constraints/conditional sampling, and repository documentation and plans. No matching report or active PR was found. Nearby #484/#512 concern datetime formatting; #414 concerns constraint discriminators; #773 concerns workflow-level repeat-until stages.
Additional context
A proposed fix direction is to apply configured numeric rounding before checking constraints and before downstream samplers evaluate conditions, consistently on each rejection-sampling attempt. Datetime/string representation postprocessing should remain separate.
Regression coverage should include scalar and column inequalities, numeric conditional dependencies, and constraints that become unsatisfiable after rounding.
I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.
Checklist
- I reproduced this issue or provided a minimal example
- I searched the docs/issues myself, or had my agent do so
- If I used an agent, I included its diagnostics above
- Dominant language
- Python
- Stars
- 2.3k
- Forks
- 219
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 42
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA-NeMo/DataDesigner
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA-NeMo/DataDesigner#995 ·
Maintainers usually reply within 1 day
-
enforce \from future import annotations` via ruff FA102 rule`Possibly taken @chethanuk claimed this 27 days ago. Opentask
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 53/100
NVIDIA-NeMo/DataDesigner#996 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 36/100
NVIDIA-NeMo/DataDesigner#994 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA-NeMo/DataDesigner#990 ·
Maintainers usually reply within 1 day
All issues in NVIDIA-NeMo/DataDesigner
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Graphify-Labs/graphify#4241 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
-
DeviceTrackerOpen
Difficulty 2/5 1-3 hours Newbie friendliness 63/100
XiaoMi/ha_xiaomi_home#1821 ·
Maintainers usually reply within 1 day
-
Maven path-index: "Ambiguous or noncanonical artifact path" error does not report the offending pathOpen
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
pulp/pulp_maven#524 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day