type: complex variable fails in for_each_task.inputs -- auto-serialize needed
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 68/100
Línea de trabajo
Start in libs/dyn/convert/ and trace how dyn.Value sequences and maps are converted to string fields, using libs/dyn/dynvar/ and bundle/config/ to understand variable resolution. Reproduce the issue with bundle validate, then verify that complex values used as for_each_task.inputs validate as JSON strings while object-typed fields remain unchanged.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
type: complex variables resolve to Go slices/maps in the config tree.
When substituted into for_each_task.inputs (which expects a JSON string),
the DABs CLI rejects the type mismatch locally -- the Jobs API never
sees the value. The fix is a json.Marshal() coercion in the CLI's
variable resolution pipeline.
Reproduction
variables:
table_list:
type: complex
default:
- task_id: load_customers
table: customers
source_type: parquet
- task_id: load_orders
table: orders
source_type: json
- task_id: load_products
table: products
source_type: csv
resources:
jobs:
my_job:
tasks:
- task_key: process_tables
for_each_task:
inputs: ${var.table_list} # "expected string, found sequence"
bundle validate fails. Both inputs: ${var.table_list} and
inputs: "${var.table_list}" produce the same error.
Per the DABs documentation,
complex variables work in fields that expect objects, e.g.:
# From Databricks docs -- new_cluster accepts an object, so complex var works
new_cluster: ${var.my_cluster}
The difference: new_cluster expects an object (complex var fits directly),
for_each_task.inputs expects a JSON string (complex var is a sequence -- type mismatch).
Root cause
The error occurs in the DABs CLI, not in the Jobs API:
1. CLI reads YAML config
2. CLI resolves ${var.table_list}
--> inserts Go []interface{} into the dyn.Value config tree
3. CLI validates config tree against Jobs API schema
4. Schema says: for_each_task.inputs --> type: string
5. Resolved value is: dyn.KindSequence ([]interface{})
6. Type mismatch --> "expected string, found sequence"
7. Jobs API is NEVER called
Relevant code locations in github.com/databricks/cli:
libs/dyn/dynvar/-- resolves${var.xxx}references, substitutes
resolveddyn.Value(KindSequence / KindMap for complex variables)
into the config treelibs/dyn/convert/-- converts the dyn.Value tree to typed Go
structs; this is where string vs. sequence mismatch is caught and
the error is raisedbundle/config/-- merges variable defaults, target overrides,
and resolved references into the final config
Proposed fix
In the conversion layer (libs/dyn/convert/), when a dyn.Value of
KindSequence or KindMap is assigned to a field typed as string:
// Pseudocode -- in the type conversion path
case dyn.KindSequence, dyn.KindMap:
if targetField.Type == reflect.String {
// Auto-serialize complex value to JSON string
jsonBytes, err := json.Marshal(value.AsAny())
if err != nil {
return dyn.InvalidValue, fmt.Errorf("cannot serialize complex variable to string: %w", err)
}
return dyn.NewValue(string(jsonBytes), value.Locations()), nil
}
This is a localized change -- one coercion rule in the existing type
conversion pipeline:
- No new syntax
- No new language concepts (DABs has no functions;
${var.xxx}is pure property access) - No Jobs API changes
- No schema changes
- The CLI already knows both the source type (complex variable = sequence/map)
and the target type (string, from the schema). It just needs to bridge the
gap withjson.Marshal().
Why this is safe: for_each_task.inputs is documented as a JSON string
encoding an array of objects. Serializing a sequence to a JSON string
produces exactly what the user would write manually:
# What users write today (manual JSON string):
inputs: >-
[{"task_id":"load_customers","table":"customers"},{"task_id":"load_orders","table":"orders"}]
# What the fix enables (CLI does the serialization):
inputs: ${var.table_list} # complex var --> json.Marshal() --> identical JSON string
The deployed job definition is byte-identical in both cases.
Expected behavior after fix
# Config: readable, validated, native YAML
variables:
bronze_ingestion_inputs:
description: "Tables to load from source"
type: complex
default:
- task_id: ingest.customers
table: customers
source_type: parquet
load_type: incremental
- task_id: ingest.orders
table: orders
source_type: json
load_type: incremental
- task_id: ingest.products
table: products
source_type: csv
load_type: full
# Job: fully declarative, no bridge notebook
resources:
jobs:
daily_etl:
name: "Daily ETL Pipeline"
tasks:
- task_key: bronze_ingestion
for_each_task:
inputs: ${var.bronze_ingestion_inputs} # just works
concurrency: 10
task:
task_key: load_table
notebook_task:
notebook_path: ./notebooks/load_to_bronze
base_parameters:
table: "{{input.table}}"
source_type: "{{input.source_type}}"
load_type: "{{input.load_type}}"
Current workaround cost
Without this fix, users must create a bridge notebook that reads config
files from workspace disk, serializes them to JSON, and passes them via
dbutils.jobs.taskValues.set():
- Extra task in every job run (~3-5 sec overhead)
- ~60 lines of Python to maintain
- Runtime dependency on workspace file I/O
- Breaks the declarative pattern -- job YAML references task values
instead of variables
At scale (50-200 tables, 5-10 for_each_task blocks per job, multiple
jobs per region), this bridge notebook becomes a runtime dependency on
every single job execution.
Environment
- Databricks on Azure
- DABs CLI (latest)
- Tested with
bundle validateandbundle deploy
- Lenguaje dominante
- Go
- Estrellas
- 403
- Forks
- 244
- Merge medio
- 1 d 13 h
- PR fusionados (30 d)
- 263
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de databricks/cli
-
CLI
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
databricks/cli#6910 ·
Los mantenedores suelen responder en 1 día
-
DABs
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
databricks/cli#6670 ·
Los mantenedores suelen responder en 1 día
-
DABs PyDABs
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
databricks/cli#3926 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 68/100
databricks/cli#6786 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
databricks/cli#6785 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de databricks/cli
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día
-
kind/engineering pulumi/pulumi-terraform Task Workflow Failure
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
pulumi/pulumi-terraform#1215 ·
Los mantenedores suelen responder en 1 día
-
bug needs-acceptance wg/developer-experience-ecosystem
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
vllm-project/semantic-router#4480 ·
Los mantenedores suelen responder en 1 día
-
area/docs kind/documentation priority/backlog triage/accepted
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
lexfrei/cloudflare-tunnel-gateway-controller#943 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
keyxmakerx/Chronicle#967 ·
Los mantenedores suelen responder en 1 día