Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Production-scale import for multi-GB CSV, JSON, and Parquet files

未關閉
#240 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 2 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
25/100
Issue 類型
功能
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
azure
領域
cli, databases

研究方向

Start by reading the existing import command and its parsing and write paths; the issue does not name specific files or tests. Compare the current streaming and serial-write behavior with the acceptance criteria, then identify the existing test entry points for import formats and error handling. Done means the documented formats, bounded concurrency, checkpoint/resume, progress, and failure output meet the criteria while existing JSON and CSV behavior remains compatible.

由索引模型根據 Issue 內容生成。

描述

enhancement

Customer scenario

Importing a large local dataset (for example, a 10 GB CSV, JSON/JSONL, or Parquet file) into Azure Cosmos DB currently often pushes customers toward Spark or Databricks. That adds significant setup, learning, governance, and pricing overhead for what should be a straightforward data-loading workflow.

The shell should provide a production-ready experience comparable in simplicity to mongoimport: point the command at a file, connection/container, and relevant options, then let it safely and efficiently load the data.

Current behavior

The existing import command provides a useful foundation:

  • JSON Lines, JSON array, and CSV input
  • Streaming parsing rather than loading the entire file into memory
  • Insert and upsert modes
  • --dry-run validation
  • --continue-on-error
  • Partition-key mapping for CSV
  • Imported/failed counts and total RU charge

However, imports currently issue item writes serially. The command does not support Parquet, checkpoint/resume, configurable concurrency, durable failure output, or detailed progress reporting. As a result, it should not yet be positioned as an optimized or resilient multi-GB migration tool.

Proposed behavior

Extend import into a production-scale data loader while preserving its simple defaults.

Suggested capabilities:

  • Add Parquet input support.
  • Use bounded, configurable parallelism and Cosmos DB bulk execution where appropriate.
  • Respect throttling and retry guidance while allowing an optional RU or throughput budget.
  • Show periodic progress: records processed, succeeded, failed, elapsed time, throughput, RU consumed, and estimated completion when available.
  • Support checkpoints and resuming an interrupted import without starting from the beginning.
  • Write rejected records and structured error details to a user-selected file.
  • Support common field mapping and type-conversion needs, especially for CSV and Parquet.
  • Preserve streaming behavior and bounded memory usage for all formats.
  • Retain insert/upsert, dry-run, and continue-on-error behavior.
  • Produce both human-readable and structured summaries suitable for automation.

Possible usage:

import ./items.jsonl --mode=upsert --max-concurrency=16 --checkpoint=./items.checkpoint
import ./items.csv --partition-key=/tenantId --errors=./rejected.jsonl
import ./items.parquet --format=parquet --max-ru=50000

The exact option names are illustrative and should follow existing shell conventions.

Acceptance criteria

  • A multi-GB input file is processed with bounded memory usage.
  • JSONL, JSON array, and CSV behavior remains backward compatible.
  • Parquet files can be imported with documented type-mapping behavior.
  • Users can configure bounded write concurrency.
  • The importer handles service throttling without losing or silently skipping records.
  • An interrupted import can resume from a durable checkpoint.
  • Progress and final summaries include processed/succeeded/failed counts and observed RU charge.
  • Failed records can be written to a machine-readable file with actionable error details.
  • Cancellation leaves a valid checkpoint and does not report success.
  • Documentation clearly explains performance, consistency, idempotency, partition-key, and resume semantics.

Notes

This is intended for scalable client-side ingestion, not a transactional import. Writes across logical partitions cannot be made atomic as a single operation.

主要語言
C#
星號
4
分支
7
平均合併
2 天 11 分鐘
30 天內合併 PR
28

環境準備

  • 沒有 Dockerfile 或 Docker Compose 檔案
  • 沒有 Pull Request 範本
  • 閱讀貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

Azure/CosmosDBShell 的其他 Issue

查看 Azure/CosmosDBShell 的全部 Issue

相似的 Issue

更多 C# Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。