Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Embedding column crashes on the stringified JSON list its docstring documents

Open
#994 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
36/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
data

Research direction

Start with _prepare_embedding_inputs in packages/data-designer-engine/src/data_designer/engine/column_generators/generators/embedding.py and the cited deserialize_json_values and parse_list_string behavior in processing/utils.py. Read the embedding tests at test_embedding.py and test_async_generators.py, then run the engine package tests. Done means the documented JSON-list input yields one embedding per element and valid-JSON non-list text is embedded as written; a fix is already prepared, so coordinate with the reporter before starting.

Written by the indexing model from the issue text.

Description

Priority Level

High (Major functionality broken)

Describe the bug

An embedding column crashes on the one input format its own documentation names.

EmbeddingColumnConfig.target_column says:

The column could be a single text string or a list of text strings in stringified JSON format. If it is a list of text strings in stringified JSON format, the embeddings will be generated for each text string.

packages/data-designer-config/src/data_designer/config/column_configs.py:571.

_prepare_embedding_inputs does this:

def _prepare_embedding_inputs(self, data: dict) -> list[str]:
    deserialized_record = deserialize_json_values(data)
    return parse_list_string(deserialized_record[self.config.target_column])

packages/data-designer-engine/src/data_designer/engine/column_generators/generators/embedding.py:30.

deserialize_json_values has already turned the stringified JSON into a real Python list, and parse_list_string opens with text = text.strip() (processing/utils.py:118), so it gets a list and raises. The documented format never works.

The same applies to any cell whose text happens to be valid JSON. A text column holding 123, true or null arrives at parse_list_string as int, bool or None and raises the same way. That reaches an LLM text column that returns a bare number and a sampler column with convert_to="str".

What does work is the Python repr with single quotes, "['a', 'b']", which is not valid JSON, so deserialize_json_values leaves it a string and parse_list_string handles it. That undocumented form is what both existing tests use (test_embedding.py:42 and test_async_generators.py:418), which is why nothing catches this.

Steps/Code to reproduce bug

On main at 0dd9f0f, Python 3.12.12, with the repo's own stub_resource_provider fixture so no model or API key is needed:

from unittest.mock import patch

import pytest

from data_designer.config.column_configs import EmbeddingColumnConfig
from data_designer.engine.column_generators.generators.embedding import EmbeddingCellGenerator


@pytest.mark.parametrize(
    "value,expected",
    [
        ('["hello", "world"]', ["hello", "world"]),   # documented stringified JSON list
        ("['hello', 'world']", ["hello", "world"]),   # undocumented python repr
        ("hello world", ["hello world"]),
        ("123", ["123"]),
        ("true", ["true"]),
    ],
)
def test_embedding_inputs(stub_resource_provider, value, expected):
    config = EmbeddingColumnConfig(name="emb", target_column="text", model_alias="test_model")
    gen = EmbeddingCellGenerator(config=config, resource_provider=stub_resource_provider)
    with patch.object(
        stub_resource_provider.model_registry.get_model.return_value,
        "generate_text_embeddings",
        return_value=[[0.1]],
    ) as mock_generate:
        gen.generate(data={"text": value})
    mock_generate.assert_called_once_with(input_texts=expected)

3 failed, 2 passed. The two that pass are the python repr and the plain string. The three that fail all end the same way:

src/data_designer/engine/column_generators/generators/embedding.py:32: in _prepare_embedding_inputs
    return parse_list_string(deserialized_record[self.config.target_column])
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _

text = True

    def parse_list_string(text: str) -> list[str]:
        """Parse a list from a string, handling JSON arrays, Python lists, and trailing commas."""
>       text = text.strip()
E       AttributeError: 'bool' object has no attribute 'strip'

src/data_designer/engine/processing/utils.py:118: AttributeError

with 'list' object has no attribute 'strip' for the JSON array and 'int' for 123.

Expected behavior

The documented stringified JSON list produces one embedding per element. A cell whose text is valid JSON but not a list produces one embedding of that text as written, so "true" embeds the four characters the user put in the column rather than crashing or becoming "True".

Agent Diagnostic / Prior Investigation

Searched for duplicates before filing: gh issue list --repo NVIDIA-NeMo/DataDesigner --state all --search embedding (nothing on this; #40 and #110 are the closed feature requests that added the column), a repo-wide search for parse_list_string in issues (nothing), and the 6 open PRs (#993, #991, #974, #959, #943, #940; none touch embedding.py or processing/utils.py).

Baseline test runs, per package, because pytest packages from the workspace root fails collection under the pinned pytest 9 (the three tests/conftest.py files become non-top-level and their pytest_plugins is rejected): data-designer-config 659 passed, data-designer-engine 2346 passed, data-designer 1161 passed, 1 skipped, 2 failed. The two interface failures are environmental and unrelated: test_create_run_config_does_not_replace_required_dataset_config expects CONFIG_SOURCE in the usage line where the installed typer 0.27.3 prints {config_source}, and test_import_performance shells out to make perf-import through uv run.

Also worth noting while you are in that file: parse_list_string is only wrong here because of the caller, but on its own it mishandles JSON that is not an array. parse_list_string('"hello"') returns ['h', 'e', 'l', 'l', 'o'], parse_list_string('{"a": 1}') returns ['a'], and '123', 'true', 'null' raise TypeError: 'int' object is not iterable out of _clean_whitespace, which except json.JSONDecodeError does not catch. Its docstring promises "If all else fails, return the original text", and that line is unreachable for those inputs. None of this is reachable through the embedding path today, since deserialize_json_values consumes every valid-JSON string before parse_list_string sees it, so it is a latent problem rather than a user-visible one.

Additional context

I have a fix ready and would like to open the PR once this is triaged. It keeps the deserialization and then branches on what came out: a list is used element by element, a string goes through parse_list_string as before, and anything else embeds the original cell text rather than its Python repr. That makes all five cases above pass, and data-designer-engine goes from 2346 to 2351 passed with the five new parameters, nothing else changing.

Happy to split the parse_list_string hardening into its own issue and PR if you would rather keep them apart.

Dominant language
Python
Stars
2.3k
Forks
219
Avg merge
1d 16h
Merged PRs (30d)
42

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA-NeMo/DataDesigner

All issues in NVIDIA-NeMo/DataDesigner

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.