Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

datalake_fdw: timestamp columns in units other than microseconds

オープン
#1,990 コメント 0 件 リアクション 0 件 担当者 1 名 GitHub で見る

@MisterRaindrop がすでに取り組んでいます。

2026年9月13日 から。

評価

この issue はまだ評価されていません。

説明

datalake
Summary

The Parquet reader in contrib/datalake_fdw (#1951) reads only microsecond timestamps (tsu:), which is what Iceberg defines and what this module writes. Files written by other systems carry other units, and today every one of them is refused with a column stored as Arrow type "tsm:..." cannot be read as timestamp.

Cases
  • Millisecond columns (TIMESTAMP_MILLIS): common in Parquet written by Spark with spark.sql.parquet.outputTimestampType=TIMESTAMP_MILLIS, by Hive, and by many ETL tools. Multiplying by 1000 loses nothing. The question is whether a lake table should read a file whose type is not the table's type; Iceberg's spec says data files carry the table's types, so accepting them is a lenience, not a requirement.
  • Nanosecond columns (TIMESTAMP_NANOS, Iceberg v3 timestamp_ns): dividing by 1000 truncates. Refuse, or truncate and say so.
  • INT96 is already handled by coercing to microseconds. One caveat, Arrow's rather than ours (reproduced with pyarrow 21 and the same setting): Arrow's microsecond conversion assumes the nanos-of-day half is non-negative, which Spark/Hive/Impala guarantee; pyarrow's deprecated INT96 writer stores a negative one for instants before 1970 and those read wrong. Coercing to nanoseconds instead would fix that one case and break every date outside 1677..2262, including the 9999-12-31 sentinels warehouses keep. Worth an upstream report.
Where

format/arrow_decode.c: dl_arrow_decode_check() decides what a timestamp column accepts, dl_arrow_decode_value() converts. The TIMESTAMP/TIMESTAMPTZ case already distinguishes zoned from unzoned by whether the format string names a zone.

Deferred from #1951 on purpose.

主要言語
C
スター
1.4k
フォーク
248
平均マージ
4日 10時間
マージ済み PR(30日)
40

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/cloudberry のほかの issue

apache/cloudberry の issue をすべて見る

似ている issue

C の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。