Embedding column crashes on the stringified JSON list its docstring documents
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 36/100
Research direction
Start with _prepare_embedding_inputs in packages/data-designer-engine/src/data_designer/engine/column_generators/generators/embedding.py and the cited deserialize_json_values and parse_list_string behavior in processing/utils.py. Read the embedding tests at test_embedding.py and test_async_generators.py, then run the engine package tests. Done means the documented JSON-list input yields one embedding per element and valid-JSON non-list text is embedded as written; a fix is already prepared, so coordinate with the reporter before starting.
Written by the indexing model from the issue text.
Description
Priority Level
High (Major functionality broken)
Describe the bug
An embedding column crashes on the one input format its own documentation names.
EmbeddingColumnConfig.target_column says:
The column could be a single text string or a list of text strings in stringified JSON format. If it is a list of text strings in stringified JSON format, the embeddings will be generated for each text string.
packages/data-designer-config/src/data_designer/config/column_configs.py:571.
_prepare_embedding_inputs does this:
def _prepare_embedding_inputs(self, data: dict) -> list[str]:
deserialized_record = deserialize_json_values(data)
return parse_list_string(deserialized_record[self.config.target_column])
packages/data-designer-engine/src/data_designer/engine/column_generators/generators/embedding.py:30.
deserialize_json_values has already turned the stringified JSON into a real Python list, and parse_list_string opens with text = text.strip() (processing/utils.py:118), so it gets a list and raises. The documented format never works.
The same applies to any cell whose text happens to be valid JSON. A text column holding 123, true or null arrives at parse_list_string as int, bool or None and raises the same way. That reaches an LLM text column that returns a bare number and a sampler column with convert_to="str".
What does work is the Python repr with single quotes, "['a', 'b']", which is not valid JSON, so deserialize_json_values leaves it a string and parse_list_string handles it. That undocumented form is what both existing tests use (test_embedding.py:42 and test_async_generators.py:418), which is why nothing catches this.
Steps/Code to reproduce bug
On main at 0dd9f0f, Python 3.12.12, with the repo's own stub_resource_provider fixture so no model or API key is needed:
from unittest.mock import patch
import pytest
from data_designer.config.column_configs import EmbeddingColumnConfig
from data_designer.engine.column_generators.generators.embedding import EmbeddingCellGenerator
@pytest.mark.parametrize(
"value,expected",
[
('["hello", "world"]', ["hello", "world"]), # documented stringified JSON list
("['hello', 'world']", ["hello", "world"]), # undocumented python repr
("hello world", ["hello world"]),
("123", ["123"]),
("true", ["true"]),
],
)
def test_embedding_inputs(stub_resource_provider, value, expected):
config = EmbeddingColumnConfig(name="emb", target_column="text", model_alias="test_model")
gen = EmbeddingCellGenerator(config=config, resource_provider=stub_resource_provider)
with patch.object(
stub_resource_provider.model_registry.get_model.return_value,
"generate_text_embeddings",
return_value=[[0.1]],
) as mock_generate:
gen.generate(data={"text": value})
mock_generate.assert_called_once_with(input_texts=expected)
3 failed, 2 passed. The two that pass are the python repr and the plain string. The three that fail all end the same way:
src/data_designer/engine/column_generators/generators/embedding.py:32: in _prepare_embedding_inputs
return parse_list_string(deserialized_record[self.config.target_column])
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
text = True
def parse_list_string(text: str) -> list[str]:
"""Parse a list from a string, handling JSON arrays, Python lists, and trailing commas."""
> text = text.strip()
E AttributeError: 'bool' object has no attribute 'strip'
src/data_designer/engine/processing/utils.py:118: AttributeError
with 'list' object has no attribute 'strip' for the JSON array and 'int' for 123.
Expected behavior
The documented stringified JSON list produces one embedding per element. A cell whose text is valid JSON but not a list produces one embedding of that text as written, so "true" embeds the four characters the user put in the column rather than crashing or becoming "True".
Agent Diagnostic / Prior Investigation
Searched for duplicates before filing: gh issue list --repo NVIDIA-NeMo/DataDesigner --state all --search embedding (nothing on this; #40 and #110 are the closed feature requests that added the column), a repo-wide search for parse_list_string in issues (nothing), and the 6 open PRs (#993, #991, #974, #959, #943, #940; none touch embedding.py or processing/utils.py).
Baseline test runs, per package, because pytest packages from the workspace root fails collection under the pinned pytest 9 (the three tests/conftest.py files become non-top-level and their pytest_plugins is rejected): data-designer-config 659 passed, data-designer-engine 2346 passed, data-designer 1161 passed, 1 skipped, 2 failed. The two interface failures are environmental and unrelated: test_create_run_config_does_not_replace_required_dataset_config expects CONFIG_SOURCE in the usage line where the installed typer 0.27.3 prints {config_source}, and test_import_performance shells out to make perf-import through uv run.
Also worth noting while you are in that file: parse_list_string is only wrong here because of the caller, but on its own it mishandles JSON that is not an array. parse_list_string('"hello"') returns ['h', 'e', 'l', 'l', 'o'], parse_list_string('{"a": 1}') returns ['a'], and '123', 'true', 'null' raise TypeError: 'int' object is not iterable out of _clean_whitespace, which except json.JSONDecodeError does not catch. Its docstring promises "If all else fails, return the original text", and that line is unreachable for those inputs. None of this is reachable through the embedding path today, since deserialize_json_values consumes every valid-JSON string before parse_list_string sees it, so it is a latent problem rather than a user-visible one.
Additional context
I have a fix ready and would like to open the PR once this is triaged. It keeps the deserialization and then branches on what came out: a list is used element by element, a string goes through parse_list_string as before, and anything else embeds the original cell text rather than its Python repr. That makes all five cases above pass, and data-designer-engine goes from 2346 to 2351 passed with the five new parameters, nothing else changing.
Happy to split the parse_list_string hardening into its own issue and PR if you would rather keep them apart.
- Dominant language
- Python
- Stars
- 2.3k
- Forks
- 219
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 42
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA-NeMo/DataDesigner
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA-NeMo/DataDesigner#995 ·
Maintainers usually reply within 1 day
-
enforce \from future import annotations` via ruff FA102 rule`Possibly taken @chethanuk claimed this 27 days ago. Opentask
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 63/100
NVIDIA-NeMo/DataDesigner#1001 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 53/100
NVIDIA-NeMo/DataDesigner#996 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA-NeMo/DataDesigner#990 ·
Maintainers usually reply within 1 day
All issues in NVIDIA-NeMo/DataDesigner
Similar issues
-
Maven path-index: "Ambiguous or noncanonical artifact path" error does not report the offending pathOpen
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
pulp/pulp_maven#524 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 3 days
-
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
infinispan/langchain-infinispan#34 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
521xueweihan/HelloGitHub#3891 ·