Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

type: complex variable fails in for_each_task.inputs -- auto-serialize needed

Aperta
#6,901 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
3/5
Tempo stimato
1-2 giorni
Idoneità per principianti
68/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
go
Ambito
cli, tooling

Direzione di ricerca

Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

DABs

Summary

type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.

Reproduction

variables:
  table_list:
    type: complex
    default:
      - task_id: load_customers
        table: customers
        source_type: parquet
      - task_id: load_orders
        table: orders
        source_type: json
      - task_id: load_products
        table: products
        source_type: csv

resources:
  jobs:
    my_job:
      tasks:
        - task_key: process_tables
          for_each_task:
            inputs: ${var.table_list}       # "expected string, found sequence"

bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.

Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:

# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}

The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).

Root cause

The error occurs in the DABs CLI, not in the Jobs API:

1. CLI reads YAML config
2. CLI resolves ${var.table_list}
   --> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called

Relevant code locations in github.com/databricks/cli:

  • libs/dyn/dynvar/ -- resolves ${var.xxx} references, substitutes
    resolved dyn.Value (KindSequence / KindMap for complex variables)
    into the config tree
  • libs/dyn/convert/ -- converts the dyn.Value tree to typed Go
    structs; this is where string vs. sequence mismatch is caught and
    the error is raised
  • bundle/config/ -- merges variable defaults, target overrides,
    and resolved references into the final config

Proposed fix

In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:

// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
    if targetField.Type == reflect.String {
        // Auto-serialize complex value to JSON string
        jsonBytes, err := json.Marshal(value.AsAny())
        if err != nil {
            return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
        }
        return dyn.NewValue(string(jsonBytes), value.Locations()), nil
    }

This is a localized change -- one coercion rule in the existing type
conversion pipeline:

  • No new syntax
  • No new language concepts (DABs has no functions; ${var.xxx} is pure property access)
  • No Jobs API changes
  • No schema changes
  • The CLI already knows both the source type (complex variable = sequence/map)
    and the target type (string, from the schema). It just needs to bridge the
    gap with json.Marshal().

Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:

# What users write today (manual JSON string):
inputs: >-
  [{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]

# What the fix enables (CLI does the serialization):
inputs: ${var.table_list}   # complex var --> json.Marshal() --> identical JSON string

The deployed job definition is byte-identical in both cases.

Expected behavior after fix

# Config: readable, validated, native YAML
variables:
  bronze_ingestion_inputs:
    description: "Tables to load from source"
    type: complex
    default:
      - task_id: ingest.customers
        table: customers
        source_type: parquet
        load_type: incremental
      - task_id: ingest.orders
        table: orders
        source_type: json
        load_type: incremental
      - task_id: ingest.products
        table: products
        source_type: csv
        load_type: full
# Job: fully declarative, no bridge notebook
resources:
  jobs:
    daily_etl:
      name: "Daily ETL Pipeline"
      tasks:
        - task_key: bronze_ingestion
          for_each_task:
            inputs: ${var.bronze_ingestion_inputs}   # just works
            concurrency: 10
            task:
              task_key: load_table
              notebook_task:
                notebook_path: ./notebooks/load_to_bronze
                base_parameters:
                  table: "{{input.table}}"
                  source_type: "{{input.source_type}}"
                  load_type: "{{input.load_type}}"

Current workaround cost

Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():

  • Extra task in every job run (~3-5 sec overhead)
  • ~60 lines of Python to maintain
  • Runtime dependency on workspace file I/O
  • Breaks the declarative pattern -- job YAML references task values
    instead of variables

At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.

Environment

  • Databricks on Azure
  • DABs CLI (latest)
  • Tested with bundle validate and bundle deploy
Lingua principale
Go
Stelle
403
Fork
244
Merge medio
1g 13h
PR unite (30g)
263

Preparare l'ambiente

  • Include un Dockerfile o un file Docker Compose
  • Ha un modello di pull request
  • Nessuna guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di databricks/cli

Tutte le issue di databricks/cli

Issue simili

Altre issue su Go

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.