Numeric sampler rounding can violate constraints and invalidate conditional sampling
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 63/100
Línea de trabajo
Read packages/data-designer-engine/src/data_designer/engine/sampling_gen/generator.py, especially _run_rejection_sampling() and generate(), and inspect the generator and constraint tests. Reproduce the rounding cases from the issue, then add regression coverage for constraints and conditional dependencies using final rounded values. Done means returned values satisfy configured constraints and dependent conditions reflect the values in the returned dataset, including the unsatisfiable-after-rounding case.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Priority Level
Medium (Annoying but has workaround)
Describe the bug
Setting decimal_places on a numeric sampler can silently produce final values that violate an explicit ScalarInequalityConstraint. Conditional samplers can also select a branch inconsistent with the final value of their source column.
Both behaviors arise because decimal rounding happens after rejection sampling and after dependent columns have been generated. Generation succeeds, but the returned data violates its configured rules.
Steps/Code to reproduce bug
Reproduced on an unmodified checkout of main at 0b8c669c717ed4806a96f8dd3629868900045dc5, using Python 3.10.12 and dependencies installed with uv sync --all-packages --frozen. No model calls are needed.
from data_designer.config import ScalarInequalityConstraint
from data_designer.engine.sampling_gen.generator import DatasetGenerator
from data_designer.engine.sampling_gen.schema import DataSchema
schema = DataSchema(
columns=[{
"name": "value",
"sampler_type": "uniform",
"params": {"low": 0, "high": 1, "decimal_places": 0},
}],
constraints=[ScalarInequalityConstraint(
target_column="value", operator="gt", rhs=0.25,
)],
)
frame = DatasetGenerator(
sampler_columns=None, schema=schema, random_state=42,
).generate(256)
print((frame["value"] <= 0.25).sum()) # 78
assert (frame["value"] > 0.25).all() # fails
Actual: 78 of 256 final values are 0, violating value > 0.25, without an error. Changing only decimal_places to None yields zero violations.
A second reproduction demonstrates inconsistent conditional labels:
schema = DataSchema(
columns=[
{
"name": "value",
"sampler_type": "uniform",
"params": {"low": 0.41, "high": 0.49, "decimal_places": 0},
},
{
"name": "label",
"sampler_type": "category",
"params": {"values": ["low"]},
"conditional_params": {
"value > 0.4": {"values": ["high"]},
},
},
],
constraints=[],
)
frame = DatasetGenerator(
sampler_columns=None, schema=schema, random_state=42,
).generate(16)
expected = frame["value"].apply(lambda value: "high" if value > 0.4 else "low")
print(frame.head())
print((frame["label"] != expected).sum()) # 16
All 16 records contain value=0.0 and label="high", although the final source values make that condition false. Changing decimal_places to None yields no contradictory labels. This conditional case was also reproduced through DataDesigner.create() and load_dataset().
Expected behavior
Every returned numeric value satisfies the configured constraint. The first configuration is feasible: rounded values of 1 satisfy value > 0.25; values that round to 0 should be rejected and resampled.
Conditional samplers should evaluate the source values that will appear in the resulting dataset, so the second example should produce label="low".
Agent Diagnostic / Prior Investigation
Investigation and reproductions were performed with a coding agent.
In generator.py:
_run_rejection_sampling()runspreprocand then checks constraints.generate()completes sampling for all columns, including dependent conditional columns.- Only afterward does it call
_round_if_needed().
Consequently, constraints and dependent conditions see pre-rounding values that differ from the returned dataset.
Existing generator tests assert constraints and conditional relationships on final returned values, while precision tests cover decimal_places independently. The sampling documentation does not describe an exception allowing rounded output to violate constraints.
Validation against the unmodified checkout:
- Existing
test_generator.pyandtest_constraints.py: 35 passed. - Four additional behavioral checks: both unrounded controls pass; the rounded constraint and conditional cases fail (2 passed, 2 failed).
- Reproduction imports were verified to resolve to the unmodified checkout.
Duplicate check on 2026-10-08 covered all 50 open issues and 5 open PRs, the 100 most recently updated closed PRs, historical issue/PR searches for precision/rounding/constraints/conditional sampling, and repository documentation and plans. No matching report or active PR was found. Nearby #484/#512 concern datetime formatting; #414 concerns constraint discriminators; #773 concerns workflow-level repeat-until stages.
Additional context
A proposed fix direction is to apply configured numeric rounding before checking constraints and before downstream samplers evaluate conditions, consistently on each rejection-sampling attempt. Datetime/string representation postprocessing should remain separate.
Regression coverage should include scalar and column inequalities, numeric conditional dependencies, and constraints that become unsatisfiable after rounding.
I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.
Checklist
- I reproduced this issue or provided a minimal example
- I searched the docs/issues myself, or had my agent do so
- If I used an agent, I included its diagnostics above
- Lenguaje dominante
- Python
- Estrellas
- 2.3k
- Forks
- 219
- Merge medio
- 1 d 16 h
- PR fusionados (30 d)
- 42
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA-NeMo/DataDesigner
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
NVIDIA-NeMo/DataDesigner#995 ·
Los mantenedores suelen responder en 1 día
-
enforce \from future import annotations` via ruff FA102 rule`Posiblemente ocupada @chethanuk la tomó hace 28 días. Abiertotask
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 53/100
NVIDIA-NeMo/DataDesigner#996 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 36/100
NVIDIA-NeMo/DataDesigner#994 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
NVIDIA-NeMo/DataDesigner#990 ·
Los mantenedores suelen responder en 1 día
Todos los issues de NVIDIA-NeMo/DataDesigner
Issues similares
-
bug ready for review
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
odysseus-dev/odysseus#6641 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
happypawspillaro/happypaws#78 ·
Los mantenedores suelen responder en 4 días
-
pydanty:is-working
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
pydantic/pydantic-ai#10020 ·
Los mantenedores suelen responder en 1 día
-
Bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
ansible-collections/ibm_zos_core#2650 ·
-
hw: pvc tests: vllm vllm
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
intel/intel-xpu-backend-for-triton#8362 ·
Los mantenedores suelen responder en 1 día