Jinja2 templates cannot reference columns created by PRE_BATCH processors
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 48/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- backend, data-engineering
Hướng nghiên cứu
Bắt đầu với raw seed reader của compiler và phần xử lý ProcessorConfig/DropColumnsProcessor được mô tả trong issue. Theo dõi cách compiler xây dựng tập hợp các cột và xác thực các template Jinja2 cùng các dependency của DAG. Hoàn thành nghĩa là các phần bổ sung và loại bỏ trong PRE_BATCH được phản ánh trong post-processor schema, cho phép các tham chiếu downstream như {{ c }} trong khi vẫn giữ nguyên hành vi của POST_BATCH và AFTER_GENERATION.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Bug Description
Jinja2 {{ }} references in LLM column prompts fail when the referenced column is created by a PRE_BATCH processor. The compiler validates templates against the raw seed schema, which doesn't include columns added at runtime by processors.
Steps to Reproduce
- Create a workflow with a seed dataset that has columns
[a, b] - Add a PRE_BATCH processor that creates a new column
c - Define a downstream LLM column whose prompt references
{{ c }} - Run the workflow
Expected Behavior
The compiler should recognize that column c will exist after the PRE_BATCH processor runs, and allow {{ c }} in downstream prompts.
Actual Behavior
The compiler rejects the template because c doesn't exist in the raw seed data. The validation happens at compile time against the raw schema, before any processors run.
Root Cause
The compiler discovers seed columns from the raw seed reader and validates Jinja2 templates against that raw schema. PRE_BATCH processors can add columns at runtime, but the compiler has no way to know about them.
DropColumnsProcessor already implicitly declares removed columns via its config - the builder uses it to mark columns with drop=True at build time. But there's no equivalent mechanism for declaring added columns.
Proposed Fix
PRE_BATCH processors should declare which columns they add/remove so the compiler can compute the post-processor column set:
class ProcessorConfig(ConfigBase):
processor_type: str
columns_added: list[str] = []
columns_removed: list[str] = []
The compiler would adjust the column set after seed column discovery: remove declared drops, add SeedDatasetColumnConfig entries for declared additions. Template validation and DAG resolution would then see the final schema.
This only applies to PRE_BATCH processors - POST_BATCH and AFTER_GENERATION processors don't need this since no downstream generators depend on their output schema.
- Ngôn ngữ chính
- Python
- Star
- 2.3k
- Fork
- 215
- Merge trung bình
- 3 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 45
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA-NeMo/DataDesigner
-
task
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 78/100
NVIDIA-NeMo/DataDesigner#971 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Run Slurm generation without a client image using a versioned vLLM serving imageCó thể đã có người làm @nabinchha đã nhận hôm nay. Đang mởtask triaged
NVIDIA-NeMo/DataDesigner#969 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Harden Slurm inference routing, backpressure, and failoverCó thể đã có người làm @nabinchha đã nhận 1 ngày trước. Đang mởtask
NVIDIA-NeMo/DataDesigner#966 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement triaged
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 40/100
NVIDIA-NeMo/DataDesigner#956 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA-NeMo/DataDesigner
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
kornia/kornia#5263 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Metadata correction for W16-5400Đang mởapproved correction metadata
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
acl-org/acl-anthology#10133 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
BasedHardware/omi#20084 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-acceptance wg/evaluation-quality
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
vllm-project/semantic-router#4424 ·
Maintainer thường phản hồi trong vòng 1 ngày