Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

docs: required_columns description is incomplete for LLM and multimodal columns

Open Beginner friendly
#1,002 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@nightcityblade is already working on this.

Since Oct 9, 2026.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
65/100
Issue type
Documentation
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
documentation

Research direction

Start with the required_columns bullet list in fern/versions/latest/pages/concepts/columns.mdx (around lines 222-230), then read the required_columns properties in packages/data-designer-config/src/data_designer/config/column_configs.py for LLMTextColumnConfig, ImageColumnConfig and EmbeddingColumnConfig. Done when the docs list every source: prompt and system_prompt Jinja2 variables, multi_modal_context column names, and target_column for embeddings. Confirm the intended behavior with a maintainer before editing, since the issue is awaiting triage.

Written by the indexing model from the issue text.

Description

bug
Priority Level

Low (Cosmetic / Minor)

Describe the bug

The concept documentation for required_columns in fern/versions/latest/pages/concepts/columns.mdx states that for LLM/Expression columns, dependencies are derived solely from Jinja2 {{ variables }}:

  • For LLM/Expression columns: extracted from Jinja2 template {{ variables }}

However, in packages/data-designer-config/src/data_designer/config/column_configs.py, LLMTextColumnConfig (and inheriting classes) also derives dependencies from:

  1. Jinja2 variables inside system_prompt (when provided)
  2. Column names referenced in multi_modal_context (ctx.column_name)

Additionally, ImageColumnConfig derives dependencies from multi_modal_context, and EmbeddingColumnConfig depends on target_column.

The current documentation wording is incomplete regarding how DAG column dependencies are assembled. Updating this section will help users properly understand DAG generation ordering when using multimodal context and system prompts.

Steps/Code to reproduce bug
import data_designer.config as dd

col = dd.LLMTextColumnConfig(
    name="summary",
    model_alias="test-model",
    prompt="Summarize this text",
    system_prompt="Context: {{ user_persona }}",
    multi_modal_context=[dd.ImageContext(column_name="image_input")]
)

# Inspect computed required_columns
print(col.required_columns)
# Output: ['user_persona', 'image_input']
# The docs state required_columns only extracts from Jinja2 template prompt,
# missing system_prompt and multi_modal_context dependencies.
Expected behavior

The required_columns section in columns.mdx should accurately list all sources of dependencies for LLM and multimodal columns:

  • Variables in prompt and system_prompt Jinja2 templates ({{ variables }})
  • Column names referenced in multi_modal_context
  • target_column for embedding columns
Agent Diagnostic / Prior Investigation

Investigated the documentation and source codebase:

  • Checked fern/versions/latest/pages/concepts/columns.mdx (lines 222–230).
  • Inspected implementation in packages/data-designer-config/src/data_designer/config/column_configs.py:
    • LLMTextColumnConfig.required_columns (lines 197–209): extracts from self.prompt, self.system_prompt, and self.multi_modal_context.
    • ImageColumnConfig.required_columns (lines 637–648): extracts from self.prompt and self.multi_modal_context.
    • EmbeddingColumnConfig.required_columns (lines 589–591): returns [self.target_column].
  • Confirmed that dataset builders rely on these exact dependencies for topological sorting in the DAG execution plan.
Additional context

I would be happy to contribute a fix after this issue is triaged and the intended behavior is confirmed.

Checklist
  • I reproduced this issue or provided a minimal example
  • I searched the docs/issues myself, or had my agent do so
  • If I used an agent, I included its diagnostics above
Dominant language
Python
Stars
2.3k
Forks
219
Avg merge
1d 21h
Merged PRs (30d)
38

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA-NeMo/DataDesigner

All issues in NVIDIA-NeMo/DataDesigner

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.