Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

type: complex variable fails in for_each_task.inputs -- auto-serialize needed

Open
#6,901 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
68/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
go
Domain
cli, tooling

Research direction

Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.

Written by the indexing model from the issue text.

Description

DABs

Summary

type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.

Reproduction

variables:
  table_list:
    type: complex
    default:
      - task_id: load_customers
        table: customers
        source_type: parquet
      - task_id: load_orders
        table: orders
        source_type: json
      - task_id: load_products
        table: products
        source_type: csv

resources:
  jobs:
    my_job:
      tasks:
        - task_key: process_tables
          for_each_task:
            inputs: ${var.table_list}       # "expected string, found sequence"

bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.

Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:

# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}

The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).

Root cause

The error occurs in the DABs CLI, not in the Jobs API:

1. CLI reads YAML config
2. CLI resolves ${var.table_list}
   --> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called

Relevant code locations in github.com/databricks/cli:

  • libs/dyn/dynvar/ -- resolves ${var.xxx} references, substitutes
    resolved dyn.Value (KindSequence / KindMap for complex variables)
    into the config tree
  • libs/dyn/convert/ -- converts the dyn.Value tree to typed Go
    structs; this is where string vs. sequence mismatch is caught and
    the error is raised
  • bundle/config/ -- merges variable defaults, target overrides,
    and resolved references into the final config

Proposed fix

In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:

// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
    if targetField.Type == reflect.String {
        // Auto-serialize complex value to JSON string
        jsonBytes, err := json.Marshal(value.AsAny())
        if err != nil {
            return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
        }
        return dyn.NewValue(string(jsonBytes), value.Locations()), nil
    }

This is a localized change -- one coercion rule in the existing type
conversion pipeline:

  • No new syntax
  • No new language concepts (DABs has no functions; ${var.xxx} is pure property access)
  • No Jobs API changes
  • No schema changes
  • The CLI already knows both the source type (complex variable = sequence/map)
    and the target type (string, from the schema). It just needs to bridge the
    gap with json.Marshal().

Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:

# What users write today (manual JSON string):
inputs: >-
  [{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]

# What the fix enables (CLI does the serialization):
inputs: ${var.table_list}   # complex var --> json.Marshal() --> identical JSON string

The deployed job definition is byte-identical in both cases.

Expected behavior after fix

# Config: readable, validated, native YAML
variables:
  bronze_ingestion_inputs:
    description: "Tables to load from source"
    type: complex
    default:
      - task_id: ingest.customers
        table: customers
        source_type: parquet
        load_type: incremental
      - task_id: ingest.orders
        table: orders
        source_type: json
        load_type: incremental
      - task_id: ingest.products
        table: products
        source_type: csv
        load_type: full
# Job: fully declarative, no bridge notebook
resources:
  jobs:
    daily_etl:
      name: "Daily ETL Pipeline"
      tasks:
        - task_key: bronze_ingestion
          for_each_task:
            inputs: ${var.bronze_ingestion_inputs}   # just works
            concurrency: 10
            task:
              task_key: load_table
              notebook_task:
                notebook_path: ./notebooks/load_to_bronze
                base_parameters:
                  table: "{{input.table}}"
                  source_type: "{{input.source_type}}"
                  load_type: "{{input.load_type}}"

Current workaround cost

Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():

  • Extra task in every job run (~3-5 sec overhead)
  • ~60 lines of Python to maintain
  • Runtime dependency on workspace file I/O
  • Breaks the declarative pattern -- job YAML references task values
    instead of variables

At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.

Environment

  • Databricks on Azure
  • DABs CLI (latest)
  • Tested with bundle validate and bundle deploy
Dominant language
Go
Stars
403
Forks
244
Avg merge
1d 13h
Merged PRs (30d)
263

Getting set up

  • Ships a Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from databricks/cli

All issues in databricks/cli

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.