Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Numeric sampler rounding can violate constraints and invalidate conditional sampling

Abierto
#1,001 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
63/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
python
Área
data

Línea de trabajo

Read packages/data-designer-engine/src/data_designer/engine/sampling_gen/generator.py, especially _run_rejection_sampling() and generate(), and inspect the generator and constraint tests. Reproduce the rounding cases from the issue, then add regression coverage for constraints and conditional dependencies using final rounded values. Done means returned values satisfy configured constraints and dependent conditions reflect the values in the returned dataset, including the unsatisfiable-after-rounding case.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Priority Level

Medium (Annoying but has workaround)

Describe the bug

Setting decimal_places on a numeric sampler can silently produce final values that violate an explicit ScalarInequalityConstraint. Conditional samplers can also select a branch inconsistent with the final value of their source column.

Both behaviors arise because decimal rounding happens after rejection sampling and after dependent columns have been generated. Generation succeeds, but the returned data violates its configured rules.

Steps/Code to reproduce bug

Reproduced on an unmodified checkout of main at 0b8c669c717ed4806a96f8dd3629868900045dc5, using Python 3.10.12 and dependencies installed with uv sync --all-packages --frozen. No model calls are needed.

from data_designer.config import ScalarInequalityConstraint
from data_designer.engine.sampling_gen.generator import DatasetGenerator
from data_designer.engine.sampling_gen.schema import DataSchema

schema = DataSchema(
    columns=[{
        "name": "value",
        "sampler_type": "uniform",
        "params": {"low": 0, "high": 1, "decimal_places": 0},
    }],
    constraints=[ScalarInequalityConstraint(
        target_column="value", operator="gt", rhs=0.25,
    )],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(256)

print((frame["value"] <= 0.25).sum())  # 78
assert (frame["value"] > 0.25).all()  # fails

Actual: 78 of 256 final values are 0, violating value > 0.25, without an error. Changing only decimal_places to None yields zero violations.

A second reproduction demonstrates inconsistent conditional labels:

schema = DataSchema(
    columns=[
        {
            "name": "value",
            "sampler_type": "uniform",
            "params": {"low": 0.41, "high": 0.49, "decimal_places": 0},
        },
        {
            "name": "label",
            "sampler_type": "category",
            "params": {"values": ["low"]},
            "conditional_params": {
                "value > 0.4": {"values": ["high"]},
            },
        },
    ],
    constraints=[],
)
frame = DatasetGenerator(
    sampler_columns=None, schema=schema, random_state=42,
).generate(16)

expected = frame["value"].apply(lambda value: "high" if value > 0.4 else "low")
print(frame.head())
print((frame["label"] != expected).sum())  # 16

All 16 records contain value=0.0 and label="high", although the final source values make that condition false. Changing decimal_places to None yields no contradictory labels. This conditional case was also reproduced through DataDesigner.create() and load_dataset().

Expected behavior

Every returned numeric value satisfies the configured constraint. The first configuration is feasible: rounded values of 1 satisfy value > 0.25; values that round to 0 should be rejected and resampled.

Conditional samplers should evaluate the source values that will appear in the resulting dataset, so the second example should produce label="low".

Agent Diagnostic / Prior Investigation

Investigation and reproductions were performed with a coding agent.

In generator.py:

  1. _run_rejection_sampling() runs preproc and then checks constraints.
  2. generate() completes sampling for all columns, including dependent conditional columns.
  3. Only afterward does it call _round_if_needed().

Consequently, constraints and dependent conditions see pre-rounding values that differ from the returned dataset.

Existing generator tests assert constraints and conditional relationships on final returned values, while precision tests cover decimal_places independently. The sampling documentation does not describe an exception allowing rounded output to violate constraints.

Validation against the unmodified checkout:

  • Existing test_generator.py and test_constraints.py: 35 passed.
  • Four additional behavioral checks: both unrounded controls pass; the rounded constraint and conditional cases fail (2 passed, 2 failed).
  • Reproduction imports were verified to resolve to the unmodified checkout.

Duplicate check on 2026-10-08 covered all 50 open issues and 5 open PRs, the 100 most recently updated closed PRs, historical issue/PR searches for precision/rounding/constraints/conditional sampling, and repository documentation and plans. No matching report or active PR was found. Nearby #484/#512 concern datetime formatting; #414 concerns constraint discriminators; #773 concerns workflow-level repeat-until stages.

Additional context

A proposed fix direction is to apply configured numeric rounding before checking constraints and before downstream samplers evaluate conditions, consistently on each rejection-sampling attempt. Datetime/string representation postprocessing should remain separate.

Regression coverage should include scalar and column inequalities, numeric conditional dependencies, and constraints that become unsatisfiable after rounding.

I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.

Checklist
  • I reproduced this issue or provided a minimal example
  • I searched the docs/issues myself, or had my agent do so
  • If I used an agent, I included its diagnostics above
Lenguaje dominante
Python
Estrellas
2.3k
Forks
219
Merge medio
1 d 16 h
PR fusionados (30 d)
42

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de NVIDIA-NeMo/DataDesigner

Todos los issues de NVIDIA-NeMo/DataDesigner

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.