Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

type: complex variable fails in for_each_task.inputs -- auto-serialize needed

Đang mở
#6,901 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
go
Lĩnh vực
cli, tooling

Hướng nghiên cứu

Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

DABs

Summary

type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.

Reproduction

variables:
  table_list:
    type: complex
    default:
      - task_id: load_customers
        table: customers
        source_type: parquet
      - task_id: load_orders
        table: orders
        source_type: json
      - task_id: load_products
        table: products
        source_type: csv

resources:
  jobs:
    my_job:
      tasks:
        - task_key: process_tables
          for_each_task:
            inputs: ${var.table_list}       # "expected string, found sequence"

bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.

Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:

# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}

The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).

Root cause

The error occurs in the DABs CLI, not in the Jobs API:

1. CLI reads YAML config
2. CLI resolves ${var.table_list}
   --> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called

Relevant code locations in github.com/databricks/cli:

  • libs/dyn/dynvar/ -- resolves ${var.xxx} references, substitutes
    resolved dyn.Value (KindSequence / KindMap for complex variables)
    into the config tree
  • libs/dyn/convert/ -- converts the dyn.Value tree to typed Go
    structs; this is where string vs. sequence mismatch is caught and
    the error is raised
  • bundle/config/ -- merges variable defaults, target overrides,
    and resolved references into the final config

Proposed fix

In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:

// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
    if targetField.Type == reflect.String {
        // Auto-serialize complex value to JSON string
        jsonBytes, err := json.Marshal(value.AsAny())
        if err != nil {
            return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
        }
        return dyn.NewValue(string(jsonBytes), value.Locations()), nil
    }

This is a localized change -- one coercion rule in the existing type
conversion pipeline:

  • No new syntax
  • No new language concepts (DABs has no functions; ${var.xxx} is pure property access)
  • No Jobs API changes
  • No schema changes
  • The CLI already knows both the source type (complex variable = sequence/map)
    and the target type (string, from the schema). It just needs to bridge the
    gap with json.Marshal().

Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:

# What users write today (manual JSON string):
inputs: >-
  [{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]

# What the fix enables (CLI does the serialization):
inputs: ${var.table_list}   # complex var --> json.Marshal() --> identical JSON string

The deployed job definition is byte-identical in both cases.

Expected behavior after fix

# Config: readable, validated, native YAML
variables:
  bronze_ingestion_inputs:
    description: "Tables to load from source"
    type: complex
    default:
      - task_id: ingest.customers
        table: customers
        source_type: parquet
        load_type: incremental
      - task_id: ingest.orders
        table: orders
        source_type: json
        load_type: incremental
      - task_id: ingest.products
        table: products
        source_type: csv
        load_type: full
# Job: fully declarative, no bridge notebook
resources:
  jobs:
    daily_etl:
      name: "Daily ETL Pipeline"
      tasks:
        - task_key: bronze_ingestion
          for_each_task:
            inputs: ${var.bronze_ingestion_inputs}   # just works
            concurrency: 10
            task:
              task_key: load_table
              notebook_task:
                notebook_path: ./notebooks/load_to_bronze
                base_parameters:
                  table: "{{input.table}}"
                  source_type: "{{input.source_type}}"
                  load_type: "{{input.load_type}}"

Current workaround cost

Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():

  • Extra task in every job run (~3-5 sec overhead)
  • ~60 lines of Python to maintain
  • Runtime dependency on workspace file I/O
  • Breaks the declarative pattern -- job YAML references task values
    instead of variables

At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.

Environment

  • Databricks on Azure
  • DABs CLI (latest)
  • Tested with bundle validate and bundle deploy
Ngôn ngữ chính
Go
Star
403
Fork
244
Merge trung bình
1 ngày 11 giờ
Pull request đã merge (30 ngày)
276

Chuẩn bị môi trường

  • Có Dockerfile hoặc tệp Docker Compose
  • Có mẫu pull request
  • Không có hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của databricks/cli

Tất cả issue của databricks/cli

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.