Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Jinja2 templates cannot reference columns created by PRE_BATCH processors

Đang mở
#394 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
48/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Đình trệ
Công nghệ
python
Lĩnh vực
backend, data-engineering

Hướng nghiên cứu

Bắt đầu với raw seed reader của compiler và phần xử lý ProcessorConfig/DropColumnsProcessor được mô tả trong issue. Theo dõi cách compiler xây dựng tập hợp các cột và xác thực các template Jinja2 cùng các dependency của DAG. Hoàn thành nghĩa là các phần bổ sung và loại bỏ trong PRE_BATCH được phản ánh trong post-processor schema, cho phép các tham chiếu downstream như {{ c }} trong khi vẫn giữ nguyên hành vi của POST_BATCH và AFTER_GENERATION.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug triaged

Bug Description

Jinja2 {{ }} references in LLM column prompts fail when the referenced column is created by a PRE_BATCH processor. The compiler validates templates against the raw seed schema, which doesn't include columns added at runtime by processors.

Steps to Reproduce

  1. Create a workflow with a seed dataset that has columns [a, b]
  2. Add a PRE_BATCH processor that creates a new column c
  3. Define a downstream LLM column whose prompt references {{ c }}
  4. Run the workflow

Expected Behavior

The compiler should recognize that column c will exist after the PRE_BATCH processor runs, and allow {{ c }} in downstream prompts.

Actual Behavior

The compiler rejects the template because c doesn't exist in the raw seed data. The validation happens at compile time against the raw schema, before any processors run.

Root Cause

The compiler discovers seed columns from the raw seed reader and validates Jinja2 templates against that raw schema. PRE_BATCH processors can add columns at runtime, but the compiler has no way to know about them.

DropColumnsProcessor already implicitly declares removed columns via its config - the builder uses it to mark columns with drop=True at build time. But there's no equivalent mechanism for declaring added columns.

Proposed Fix

PRE_BATCH processors should declare which columns they add/remove so the compiler can compute the post-processor column set:

class ProcessorConfig(ConfigBase):
    processor_type: str
    columns_added: list[str] = []
    columns_removed: list[str] = []

The compiler would adjust the column set after seed column discovery: remove declared drops, add SeedDatasetColumnConfig entries for declared additions. Template validation and DAG resolution would then see the final schema.

This only applies to PRE_BATCH processors - POST_BATCH and AFTER_GENERATION processors don't need this since no downstream generators depend on their output schema.

Ngôn ngữ chính
Python
Star
2.3k
Fork
215
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
45

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA-NeMo/DataDesigner

Tất cả issue của NVIDIA-NeMo/DataDesigner

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.