Production-scale import for multi-GB CSV, JSON, and Parquet files
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
Direzione di ricerca
Start by reading the existing import command and its parsing and write paths; the issue does not name specific files or tests. Compare the current streaming and serial-write behavior with the acceptance criteria, then identify the existing test entry points for import formats and error handling. Done means the documented formats, bounded concurrency, checkpoint/resume, progress, and failure output meet the criteria while existing JSON and CSV behavior remains compatible.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Customer scenario
Importing a large local dataset (for example, a 10 GB CSV, JSON/JSONL, or Parquet file) into Azure Cosmos DB currently often pushes customers toward Spark or Databricks. That adds significant setup, learning, governance, and pricing overhead for what should be a straightforward data-loading workflow.
The shell should provide a production-ready experience comparable in simplicity to mongoimport: point the command at a file, connection/container, and relevant options, then let it safely and efficiently load the data.
Current behavior
The existing import command provides a useful foundation:
- JSON Lines, JSON array, and CSV input
- Streaming parsing rather than loading the entire file into memory
- Insert and upsert modes
--dry-runvalidation--continue-on-error- Partition-key mapping for CSV
- Imported/failed counts and total RU charge
However, imports currently issue item writes serially. The command does not support Parquet, checkpoint/resume, configurable concurrency, durable failure output, or detailed progress reporting. As a result, it should not yet be positioned as an optimized or resilient multi-GB migration tool.
Proposed behavior
Extend import into a production-scale data loader while preserving its simple defaults.
Suggested capabilities:
- Add Parquet input support.
- Use bounded, configurable parallelism and Cosmos DB bulk execution where appropriate.
- Respect throttling and retry guidance while allowing an optional RU or throughput budget.
- Show periodic progress: records processed, succeeded, failed, elapsed time, throughput, RU consumed, and estimated completion when available.
- Support checkpoints and resuming an interrupted import without starting from the beginning.
- Write rejected records and structured error details to a user-selected file.
- Support common field mapping and type-conversion needs, especially for CSV and Parquet.
- Preserve streaming behavior and bounded memory usage for all formats.
- Retain insert/upsert, dry-run, and continue-on-error behavior.
- Produce both human-readable and structured summaries suitable for automation.
Possible usage:
import ./items.jsonl --mode=upsert --max-concurrency=16 --checkpoint=./items.checkpoint
import ./items.csv --partition-key=/tenantId --errors=./rejected.jsonl
import ./items.parquet --format=parquet --max-ru=50000
The exact option names are illustrative and should follow existing shell conventions.
Acceptance criteria
- A multi-GB input file is processed with bounded memory usage.
- JSONL, JSON array, and CSV behavior remains backward compatible.
- Parquet files can be imported with documented type-mapping behavior.
- Users can configure bounded write concurrency.
- The importer handles service throttling without losing or silently skipping records.
- An interrupted import can resume from a durable checkpoint.
- Progress and final summaries include processed/succeeded/failed counts and observed RU charge.
- Failed records can be written to a machine-readable file with actionable error details.
- Cancellation leaves a valid checkpoint and does not report success.
- Documentation clearly explains performance, consistency, idempotency, partition-key, and resume semantics.
Notes
This is intended for scalable client-side ingestion, not a transactional import. Writes across logical partitions cannot be made atomic as a single operation.
- Lingua principale
- C#
- Stelle
- 4
- Fork
- 7
- Merge medio
- 1g 21h
- PR unite (30g)
- 30
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Azure/CosmosDBShell
-
bug P0
Difficoltà 4/5 3-5 giorni Idoneità per principianti 30/100
Azure/CosmosDBShell#244 ·
I maintainer di solito rispondono entro 2 giorni
-
automation P1
Difficoltà 5/5 Più di una settimana Idoneità per principianti 45/100
Azure/CosmosDBShell#178 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Lightweight headless CI distribution (no MCP, LSP, interactive UI)Forse di nuovo libera Una pull request per questa issue è stata chiusa senza essere unita. Apertaautomation P1
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
Azure/CosmosDBShell#175 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
agentic enhancement P0
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
Azure/CosmosDBShell#153 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 45/100
Azure/CosmosDBShell#107 · 4 commenti ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di Azure/CosmosDBShell
Issue simili
-
再現済み 要トリアージ 誤判定
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
yksr-melt/Meltype#421 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 64/100
Facepunch/sbox-public#12063 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
facioquo/stock-indicators-dotnet#2316 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
eriknihlen/OpenAC#219 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
ObsidianMC/Obsidian#548 ·
I maintainer di solito rispondono entro 1 giorno