Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Variant shredding cannot express a SQL-null row: the shred pipeline has no validity

Cerrado
#398 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
48/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
csharp

Línea de trabajo

Comienza con los puntos de entrada de Apache.Arrow.Operations.Shredding mencionados en el issue: ShredSchemaInferer.Infer, VariantShredder.Shred y ShreddedVariantArrayBuilder.Build; inspecciona cómo se representa la validez del almacenamiento de VariantArray. Ejecuta la reproducción proporcionada para .NET 8 y, después, verifica que las filas SQL-NULL sigan siendo null, mientras que los valores null de JSON presentes en las variantes sigan siendo válidos, incluido el comportamiento de inferencia para las filas null.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Describe the enhancement requested

Every entry point in Apache.Arrow.Operations.Shredding works on values, never on a column with validity, so there is no way to shred a nullable variant column:

ShredSchema ShredSchemaInferer.Infer(IEnumerable<VariantValue> values, ShredOptions options = null)
(byte[], IReadOnlyList<ShredResult>) VariantShredder.Shred(IEnumerable<VariantValue> values, ShredSchema schema)
VariantArray ShreddedVariantArrayBuilder.Build(ShredSchema schema, byte[] metadata, IReadOnlyList<ShredResult> rows, MemoryAllocator allocator = null)

Build produces a VariantArray whose storage struct carries no validity buffer, so every row is valid. A caller shredding a column that contains SQL NULLs has to pass a placeholder value for each null row and then repair the result afterwards.

This matters because the two things are not interchangeable. The parquet VariantShredding spec gives each encoding a distinct meaning: a SQL-NULL row is the optional group itself being null (Arrow: the storage struct's validity), while a present value holding a variant JSON null is value = basic type 0 / physical type 0. Collapsing them changes what IS NULL means for every consumer of the column, and the round trip stops being lossless.

Repro

Apache.Arrow.Operations 23.0.0 (the only published version), .NET 8:

using Apache.Arrow;
using Apache.Arrow.Operations.Shredding;
using Apache.Arrow.Operations.VariantJson;
using Apache.Arrow.Scalars.Variant;

static VariantValue Obj(int a) => VariantValue.FromObject(
    new Dictionary<string, VariantValue> { ["a"] = VariantValue.FromInt32(a) });

// Three rows whose MIDDLE row is meant to be SQL NULL. Nothing in the pipeline takes a mask,
// so the best a caller can do is put a placeholder there.
var values = new List<VariantValue> { Obj(1), VariantValue.Null, Obj(3) };

var schema = new ShredSchemaInferer().Infer(values, ShredOptions.Default);
var (metadata, rows) = VariantShredder.Shred(values, schema);
VariantArray array = ShreddedVariantArrayBuilder.Build(schema, metadata, rows);

Console.WriteLine($"storage NullCount = {array.StorageArray.NullCount}");
for (int i = 0; i < array.Length; i++)
    Console.WriteLine($"  row {i}: IsNull={array.IsNull(i),-5} logical={VariantJsonWriter.ToJson(array.GetLogicalVariantValue(i), false)}");
storage NullCount = 0
  row 0: IsNull=False logical={"a":1}
  row 1: IsNull=False logical=null      <-- wanted a NULL ROW, got a present JSON null
  row 2: IsNull=False logical={"a":3}
Current workaround

Rebuild the storage struct with a validity bitmap and re-wrap it, which reaches past the public shredding API into ArrayData:

var storage = array.StorageArray.Data;
var validity = new ArrowBuffer.BitmapBuilder(array.Length);
validity.Append(true); validity.Append(false); validity.Append(true);
var patched = new VariantArray(array.VariantType, ArrowArrayFactory.BuildArray(
    new ArrayData(storage.DataType, storage.Length, nullCount: 1, storage.Offset,
                  new[] { validity.Build() }, storage.Children, storage.Dictionary)));
// row 1: IsNull=True

Two sharp edges in it: the bitmap is built from bit 0 while the ArrayData keeps storage.Offset, so it is only correct for an unsliced array; and the placeholder still travels through VariantShredder.Shred, so it has to be a value the shredder accepts for the inferred schema.

Suggested API

An overload that carries validity through, e.g.

VariantArray ShreddedVariantArrayBuilder.Build(
    ShredSchema schema, byte[] metadata, IReadOnlyList<ShredResult> rows,
    ReadOnlySpan<bool> isNull, MemoryAllocator allocator = null);

or a validity ArrowBuffer + nullCount pair if that fits the surrounding style better. Null rows would also be excluded from ShredSchemaInferer.Infer, which today has to be done by the caller filtering the sequence.

Component(s)

C#

Lenguaje dominante
C#
Estrellas
41
Forks
31
Merge medio
20 h 56 min
PR fusionados (30 d)
14

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de apache/arrow-dotnet

Todos los issues de apache/arrow-dotnet

Issues similares

Más issues de C#

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.