type: complex variable fails in for_each_task.inputs -- auto-serialize needed
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 68/100
Direzione di ricerca
Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.
Reproduction
variables:
table_list:
type: complex
default:
- task_id: load_customers
table: customers
source_type: parquet
- task_id: load_orders
table: orders
source_type: json
- task_id: load_products
table: products
source_type: csv
resources:
jobs:
my_job:
tasks:
- task_key: process_tables
for_each_task:
inputs: ${var.table_list} # "expected string, found sequence"
bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.
Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:
# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}
The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).
Root cause
The error occurs in the DABs CLI, not in the Jobs API:
1. CLI reads YAML config
2. CLI resolves ${var.table_list}
--> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called
Relevant code locations in github.com/databricks/cli:
libs/dyn/dynvar/-- resolves${var.xxx}references, substitutes
resolveddyn.Value(KindSequence / KindMap for complex variables)
into the config treelibs/dyn/convert/-- converts the dyn.Value tree to typed Go
structs; this is where string vs. sequence mismatch is caught and
the error is raisedbundle/config/-- merges variable defaults, target overrides,
and resolved references into the final config
Proposed fix
In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:
// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
if targetField.Type == reflect.String {
// Auto-serialize complex value to JSON string
jsonBytes, err := json.Marshal(value.AsAny())
if err != nil {
return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
}
return dyn.NewValue(string(jsonBytes), value.Locations()), nil
}
This is a localized change -- one coercion rule in the existing type
conversion pipeline:
- No new syntax
- No new language concepts (DABs has no functions;
${var.xxx}is pure property access) - No Jobs API changes
- No schema changes
- The CLI already knows both the source type (complex variable = sequence/map)
and the target type (string, from the schema). It just needs to bridge the
gap withjson.Marshal().
Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:
# What users write today (manual JSON string):
inputs: >-
[{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]
# What the fix enables (CLI does the serialization):
inputs: ${var.table_list} # complex var --> json.Marshal() --> identical JSON string
The deployed job definition is byte-identical in both cases.
Expected behavior after fix
# Config: readable, validated, native YAML
variables:
bronze_ingestion_inputs:
description: "Tables to load from source"
type: complex
default:
- task_id: ingest.customers
table: customers
source_type: parquet
load_type: incremental
- task_id: ingest.orders
table: orders
source_type: json
load_type: incremental
- task_id: ingest.products
table: products
source_type: csv
load_type: full
# Job: fully declarative, no bridge notebook
resources:
jobs:
daily_etl:
name: "Daily ETL Pipeline"
tasks:
- task_key: bronze_ingestion
for_each_task:
inputs: ${var.bronze_ingestion_inputs} # just works
concurrency: 10
task:
task_key: load_table
notebook_task:
notebook_path: ./notebooks/load_to_bronze
base_parameters:
table: "{{input.table}}"
source_type: "{{input.source_type}}"
load_type: "{{input.load_type}}"
Current workaround cost
Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():
- Extra task in every job run (~3-5 sec overhead)
- ~60 lines of Python to maintain
- Runtime dependency on workspace file I/O
- Breaks the declarative pattern -- job YAML references task values
instead of variables
At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.
Environment
- Databricks on Azure
- DABs CLI (latest)
- Tested with
bundle validateandbundle deploy
- Lingua principale
- Go
- Stelle
- 403
- Fork
- 244
- Merge medio
- 1g 13h
- PR unite (30g)
- 263
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Nessuna guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di databricks/cli
-
CLI
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
databricks/cli#6910 ·
I maintainer di solito rispondono entro 1 giorno
-
DABs
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
databricks/cli#6670 ·
I maintainer di solito rispondono entro 1 giorno
-
DABs PyDABs
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
databricks/cli#3926 · 4 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 68/100
databricks/cli#6786 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
databricks/cli#6785 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di databricks/cli
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
bug needs-acceptance wg/developer-experience-ecosystem
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
vllm-project/semantic-router#4480 ·
I maintainer di solito rispondono entro 1 giorno
-
area/docs kind/documentation priority/backlog triage/accepted
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
lexfrei/cloudflare-tunnel-gateway-controller#943 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
keyxmakerx/Chronicle#967 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
DaoCloud/DaoCloud-docs#7446 ·
I maintainer di solito rispondono entro 1 giorno