type: complex variable fails in for_each_task.inputs -- auto-serialize needed
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 68/100
Research direction
Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.
Written by the indexing model from the issue text.
Description
Summary
type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.
Reproduction
variables:
table_list:
type: complex
default:
- task_id: load_customers
table: customers
source_type: parquet
- task_id: load_orders
table: orders
source_type: json
- task_id: load_products
table: products
source_type: csv
resources:
jobs:
my_job:
tasks:
- task_key: process_tables
for_each_task:
inputs: ${var.table_list} # "expected string, found sequence"
bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.
Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:
# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}
The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).
Root cause
The error occurs in the DABs CLI, not in the Jobs API:
1. CLI reads YAML config
2. CLI resolves ${var.table_list}
--> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called
Relevant code locations in github.com/databricks/cli:
libs/dyn/dynvar/-- resolves${var.xxx}references, substitutes
resolveddyn.Value(KindSequence / KindMap for complex variables)
into the config treelibs/dyn/convert/-- converts the dyn.Value tree to typed Go
structs; this is where string vs. sequence mismatch is caught and
the error is raisedbundle/config/-- merges variable defaults, target overrides,
and resolved references into the final config
Proposed fix
In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:
// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
if targetField.Type == reflect.String {
// Auto-serialize complex value to JSON string
jsonBytes, err := json.Marshal(value.AsAny())
if err != nil {
return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
}
return dyn.NewValue(string(jsonBytes), value.Locations()), nil
}
This is a localized change -- one coercion rule in the existing type
conversion pipeline:
- No new syntax
- No new language concepts (DABs has no functions;
${var.xxx}is pure property access) - No Jobs API changes
- No schema changes
- The CLI already knows both the source type (complex variable = sequence/map)
and the target type (string, from the schema). It just needs to bridge the
gap withjson.Marshal().
Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:
# What users write today (manual JSON string):
inputs: >-
[{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]
# What the fix enables (CLI does the serialization):
inputs: ${var.table_list} # complex var --> json.Marshal() --> identical JSON string
The deployed job definition is byte-identical in both cases.
Expected behavior after fix
# Config: readable, validated, native YAML
variables:
bronze_ingestion_inputs:
description: "Tables to load from source"
type: complex
default:
- task_id: ingest.customers
table: customers
source_type: parquet
load_type: incremental
- task_id: ingest.orders
table: orders
source_type: json
load_type: incremental
- task_id: ingest.products
table: products
source_type: csv
load_type: full
# Job: fully declarative, no bridge notebook
resources:
jobs:
daily_etl:
name: "Daily ETL Pipeline"
tasks:
- task_key: bronze_ingestion
for_each_task:
inputs: ${var.bronze_ingestion_inputs} # just works
concurrency: 10
task:
task_key: load_table
notebook_task:
notebook_path: ./notebooks/load_to_bronze
base_parameters:
table: "{{input.table}}"
source_type: "{{input.source_type}}"
load_type: "{{input.load_type}}"
Current workaround cost
Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():
- Extra task in every job run (~3-5 sec overhead)
- ~60 lines of Python to maintain
- Runtime dependency on workspace file I/O
- Breaks the declarative pattern -- job YAML references task values
instead of variables
At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.
Environment
- Databricks on Azure
- DABs CLI (latest)
- Tested with
bundle validateandbundle deploy
- Dominant language
- Go
- Stars
- 403
- Forks
- 244
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 263
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from databricks/cli
-
Apps commands only accept relative path to app.ymlPossibly taken A pull request linked to this issue is open or already merged. OpenCLI
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
databricks/cli#6910 ·
Maintainers usually reply within 1 day
-
DABs
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
databricks/cli#6670 ·
Maintainers usually reply within 1 day
-
DABs PyDABs
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
databricks/cli#3926 · 4 comments ·
Maintainers usually reply within 1 day
-
Feature request: allow bundle init --config-file to read from Unity Catalog VolumesPossibly taken @leprpht claimed this 8 days ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 68/100
databricks/cli#6786 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
databricks/cli#6785 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
appbaseio/reactivesearch-api#402 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
router-for-me/CLIProxyAPI#6399 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
-
settings.py flaps between reconciles: needsMigrationSetting depends on map iteration orderPossibly taken @fontaineajulien claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
pulp/pulp-operator#1691 ·