Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Production-scale import for multi-GB CSV, JSON, and Parquet files

Aperta
#240 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 2 giorni

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
azure
Ambito
cli, databases

Direzione di ricerca

Start by reading the existing import command and its parsing and write paths; the issue does not name specific files or tests. Compare the current streaming and serial-write behavior with the acceptance criteria, then identify the existing test entry points for import formats and error handling. Done means the documented formats, bounded concurrency, checkpoint/resume, progress, and failure output meet the criteria while existing JSON and CSV behavior remains compatible.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

enhancement

Customer scenario

Importing a large local dataset (for example, a 10 GB CSV, JSON/JSONL, or Parquet file) into Azure Cosmos DB currently often pushes customers toward Spark or Databricks. That adds significant setup, learning, governance, and pricing overhead for what should be a straightforward data-loading workflow.

The shell should provide a production-ready experience comparable in simplicity to mongoimport: point the command at a file, connection/container, and relevant options, then let it safely and efficiently load the data.

Current behavior

The existing import command provides a useful foundation:

  • JSON Lines, JSON array, and CSV input
  • Streaming parsing rather than loading the entire file into memory
  • Insert and upsert modes
  • --dry-run validation
  • --continue-on-error
  • Partition-key mapping for CSV
  • Imported/failed counts and total RU charge

However, imports currently issue item writes serially. The command does not support Parquet, checkpoint/resume, configurable concurrency, durable failure output, or detailed progress reporting. As a result, it should not yet be positioned as an optimized or resilient multi-GB migration tool.

Proposed behavior

Extend import into a production-scale data loader while preserving its simple defaults.

Suggested capabilities:

  • Add Parquet input support.
  • Use bounded, configurable parallelism and Cosmos DB bulk execution where appropriate.
  • Respect throttling and retry guidance while allowing an optional RU or throughput budget.
  • Show periodic progress: records processed, succeeded, failed, elapsed time, throughput, RU consumed, and estimated completion when available.
  • Support checkpoints and resuming an interrupted import without starting from the beginning.
  • Write rejected records and structured error details to a user-selected file.
  • Support common field mapping and type-conversion needs, especially for CSV and Parquet.
  • Preserve streaming behavior and bounded memory usage for all formats.
  • Retain insert/upsert, dry-run, and continue-on-error behavior.
  • Produce both human-readable and structured summaries suitable for automation.

Possible usage:

import ./items.jsonl --mode=upsert --max-concurrency=16 --checkpoint=./items.checkpoint
import ./items.csv --partition-key=/tenantId --errors=./rejected.jsonl
import ./items.parquet --format=parquet --max-ru=50000

The exact option names are illustrative and should follow existing shell conventions.

Acceptance criteria

  • A multi-GB input file is processed with bounded memory usage.
  • JSONL, JSON array, and CSV behavior remains backward compatible.
  • Parquet files can be imported with documented type-mapping behavior.
  • Users can configure bounded write concurrency.
  • The importer handles service throttling without losing or silently skipping records.
  • An interrupted import can resume from a durable checkpoint.
  • Progress and final summaries include processed/succeeded/failed counts and observed RU charge.
  • Failed records can be written to a machine-readable file with actionable error details.
  • Cancellation leaves a valid checkpoint and does not report success.
  • Documentation clearly explains performance, consistency, idempotency, partition-key, and resume semantics.

Notes

This is intended for scalable client-side ingestion, not a transactional import. Writes across logical partitions cannot be made atomic as a single operation.

Lingua principale
C#
Stelle
4
Fork
7
Merge medio
1g 21h
PR unite (30g)
30

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Azure/CosmosDBShell

Tutte le issue di Azure/CosmosDBShell

Issue simili

Altre issue su C#

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.