Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Storage] Custom Response Deserializer (response-level)

Abierto
#1,017 3 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
rust
Área
api, tooling

Línea de trabajo

Empieza con el ejemplo clients/blob_container_client.rs de los clientes generados y, después, lee los contratos DeserializeWith y Format de typespec_client_core. Traza cómo se seleccionan los formatos de respuesta y los tokens de continuación del paginador, e identifica los puntos de entrada del generador implicados. Se considera terminado cuando el cliente generado puede admitir la selección del formato a nivel de respuesta y la paginación sin editar manualmente los archivos generados.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

feature-request Storage

Currently I am working on Arrow support for List Blobs: https://github.com/Azure/azure-sdk-for-rust/pull/4796 where I have been making some hand-written edits to the generated code just to test the feasibility, and I believe I have hit the limits of current TypeSpec and Azure Core constructs.

Background

The feature being added is that we now will expose augmenting the Accept header, and passing: application/vnd.apache.arrow.stream,application/xml
should yield arrow (with XML as infallible fallback, if the server cannot for whatever reason return arrow).

The Rust emitter currently hardcodes how a response body is deserialized based on its content type (i.e. application/xml → XmlFormat). There is no way, from TypeSpec /client.tsp, to tell the emitter "for this operation, deserialize the response with my own handwritten function." This blocks any scenario where the client must decide the body format at runtime.

This is analogous to the existing field-level deserialize_with client option:

  • (i.e.. @@clientOption(BlobItem.name, "deserialize_with", "crate::models::blob_name::option::deserialize", "rust")) - but that hook is a serde field attribute that runs inside an already-chosen format, so it can't switch between two different wire formats (Arrow vs XML). We need the equivalent hook at response-level.

Therefore, a single API call can therefore return either Arrow or XML on the wire, and the client chooses the right decoder per response.

Issues

Two things, both in the emitter's generated output:

1. No response-level format / deserializer override

The emitter picks the response Format from the content type and bakes it into the return type (e.g. Response<ListBlobsResponse, XmlFormat> / Pager<ListBlobsResponse, XmlFormat>). There is no client option to substitute a custom Format marker (e.g. crate::.../ArrowOrXmlFormat) or a custom response deserializer for a given operation that has the flexibility to determine what format and deserialization path to go down at runtime.

Sub-question: how does the custom deserializer know which format it received?

However this override is designed, the runtime deserializer needs a signal to tell Arrow from XML. There are two ways to provide that signal, and they have different implications:

Option A: Body sniffing (no signature change to DeserializeWith)

The current DeserializeWith::deserialize_with(body: ResponseBody) receives only the body, not the response headers. Since Apache Arrow IPC streams begin with the ARROW1 magic bytes and XML begins with <, the deserializer can inspect the leading bytes of the body and dispatch accordingly.

  • ✅ Works with the trait exactly as it exists today - no core signature changes.
  • ✅ Robust even if the service mislabels Content-Type; it keys off the actual payload.
  • ✅ Keeps the format decision fully self-contained in the handwritten deserializer.
  • ⚠️ Relies on a magic-byte heuristic rather than the declared content type; correct for Arrow/XML but not a general "any content type" mechanism.
  • ⚠️ Requires the body to be buffered enough to peek at the header bytes (already the case here).

Option B: Pass response headers into the deserializer

Change the DeserializeWith / Format contract so the deserializer also receives the response headers (e.g. Content-Type), and dispatch on that.

  • ✅ Uses the service's declared content type - the "correct" HTTP-native signal, and generalizes to any future format negotiation.
  • ✅ No payload-shape assumptions.
  • ⚠️ Requires a breaking change to a core trait (DeserializeWith / Format) in typespec_client_core, touching every existing Format implementation and consumer.
  • ⚠️ Trusts the server's Content-Type; if it's ever wrong or missing, deserialization picks the wrong path (Arrow's fallback-to-XML makes this a real edge case to reason about).
  • ⚠️ Larger blast radius and coordination cost across the SDK.
2. Pager continuation extraction is hardcoded to XML

For paged operations, the generated pager closure extracts the continuation token with a literal xml::from_xml(&body) to read NextMarker - outside the Format / DeserializeWith path (list_blobs example):

let (status, headers, body) = rsp.deconstruct();
let res: ...Page = xml::from_xml(&body)?;   // <-- hardcoded XML, not format-driven

On an Arrow body the continuation token lives in Arrow schema metadata, not XML, so this line fails regardless of any custom Format. Even if issue 1 is solved, paging still breaks.

Current Workarounds (Illegal Edits to Generated Code)

Today the only way to make this work is to hand-edit the generated client method to sniff the content type and decode Arrow, then re-encode it back to XML so the return type (Pager<..., XmlFormat>) is satisfied. This:

  • edits files under generated/ (overwritten on the next regeneration), and
  • does an Arrow → XML → Arrow round-trip that throws away the entire performance reason for supporting Arrow.

This also opens out the discussion of how we will handle the breaking nature of this change given that XmlFormat is currently baked into the response type unfortunately 😢

Lenguaje dominante
Rust
Estrellas
8
Forks
11
Merge medio
16 h 4 min
PR fusionados (30 d)
5

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de Azure/typespec-rust

Todos los issues de Azure/typespec-rust

Issues similares

Más issues de Rust

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.