Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

Production-scale import for multi-GB CSV, JSON, and Parquet files

Aberta
#240 0 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 2 dias

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
25/100
Tipo de issue
Funcionalidade
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
azure
Domínio
cli, databases

Direção de pesquisa

Start by reading the existing import command and its parsing and write paths; the issue does not name specific files or tests. Compare the current streaming and serial-write behavior with the acceptance criteria, then identify the existing test entry points for import formats and error handling. Done means the documented formats, bounded concurrency, checkpoint/resume, progress, and failure output meet the criteria while existing JSON and CSV behavior remains compatible.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

enhancement

Customer scenario

Importing a large local dataset (for example, a 10 GB CSV, JSON/JSONL, or Parquet file) into Azure Cosmos DB currently often pushes customers toward Spark or Databricks. That adds significant setup, learning, governance, and pricing overhead for what should be a straightforward data-loading workflow.

The shell should provide a production-ready experience comparable in simplicity to mongoimport: point the command at a file, connection/container, and relevant options, then let it safely and efficiently load the data.

Current behavior

The existing import command provides a useful foundation:

  • JSON Lines, JSON array, and CSV input
  • Streaming parsing rather than loading the entire file into memory
  • Insert and upsert modes
  • --dry-run validation
  • --continue-on-error
  • Partition-key mapping for CSV
  • Imported/failed counts and total RU charge

However, imports currently issue item writes serially. The command does not support Parquet, checkpoint/resume, configurable concurrency, durable failure output, or detailed progress reporting. As a result, it should not yet be positioned as an optimized or resilient multi-GB migration tool.

Proposed behavior

Extend import into a production-scale data loader while preserving its simple defaults.

Suggested capabilities:

  • Add Parquet input support.
  • Use bounded, configurable parallelism and Cosmos DB bulk execution where appropriate.
  • Respect throttling and retry guidance while allowing an optional RU or throughput budget.
  • Show periodic progress: records processed, succeeded, failed, elapsed time, throughput, RU consumed, and estimated completion when available.
  • Support checkpoints and resuming an interrupted import without starting from the beginning.
  • Write rejected records and structured error details to a user-selected file.
  • Support common field mapping and type-conversion needs, especially for CSV and Parquet.
  • Preserve streaming behavior and bounded memory usage for all formats.
  • Retain insert/upsert, dry-run, and continue-on-error behavior.
  • Produce both human-readable and structured summaries suitable for automation.

Possible usage:

import ./items.jsonl --mode=upsert --max-concurrency=16 --checkpoint=./items.checkpoint
import ./items.csv --partition-key=/tenantId --errors=./rejected.jsonl
import ./items.parquet --format=parquet --max-ru=50000

The exact option names are illustrative and should follow existing shell conventions.

Acceptance criteria

  • A multi-GB input file is processed with bounded memory usage.
  • JSONL, JSON array, and CSV behavior remains backward compatible.
  • Parquet files can be imported with documented type-mapping behavior.
  • Users can configure bounded write concurrency.
  • The importer handles service throttling without losing or silently skipping records.
  • An interrupted import can resume from a durable checkpoint.
  • Progress and final summaries include processed/succeeded/failed counts and observed RU charge.
  • Failed records can be written to a machine-readable file with actionable error details.
  • Cancellation leaves a valid checkpoint and does not report success.
  • Documentation clearly explains performance, consistency, idempotency, partition-key, and resume semantics.

Notes

This is intended for scalable client-side ingestion, not a transactional import. Writes across logical partitions cannot be made atomic as a single operation.

Linguagem predominante
C#
Estrelas
4
Forks
7
Merge médio
2d 11min
PRs com merge (30d)
28

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de Azure/CosmosDBShell

Todas as issues de Azure/CosmosDBShell

Issues semelhantes

Mais issues de C#

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.