aws/aws-sdk-pandas

to_iceberg: conditional merge

开放

#3,173 创建于 2025年7月8日

 (3 条评论) (1 个反应) (0 位负责人)Python (630 个派生)batch import
featuregood first issuehelp wanted

仓库指标

星标
 (3,560 个星标)
PR 合并指标
 (平均合并 6天 23小时) (30 天内合并 37 个 PR)

描述

Is your feature request related to a problem? Please describe. to_iceberg method does not allow for conditional merge. This is very desired, otherwise following arguments:

    merge_cols: list[str] | None = None,
    merge_condition: Literal["update", "ignore"] = "update",

will not be able to handle non-chronological data and can overwrite more recent records.

Describe the solution you'd like Introduce one additional merge_condition literal "conditional_merge" and one optional argument conditional_merge_string.

Extend following segment of code:

    if merge_cols:
        if merge_condition == "update":
            match_condition = f"""WHEN MATCHED THEN
                UPDATE SET {", ".join([f'"{x}" = source."{x}"' for x in df.columns])}"""
        else:
            match_condition = ""

with one elif statement:

        elif merge_condition == "conditional_merge":
            match_condition = f"""WHEN MATCHED AND {conditional_merge_string} THEN
                UPDATE SET {", ".join([f'"{x}" = source."{x}"' for x in df.columns])}"""

Describe alternatives you've considered Writing Athena queries directly and bypassing entire _write_iceberg.py implementation.

贡献者指南