Allow user to define a subset of columns for update detection in UPSERT
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Anfängerfreundlichkeit
- 55/100
- Issue-Typ
- Feature
- Klarheit
- Größtenteils klar
- Aktivitätsstatus
- Ruhig
- Tech-Stack
- python
- Bereich
- data-engineering, databases
Rechercherichtung
Beginne mit der Untersuchung von get_rows_to_update und der upsert-Methoden. Vergleiche anschließend das vorgeschlagene Verhalten mit dem Commit 59d18337a61b106566c2e6a432b7d4899ca7f334. Füge eine optionale Behandlung von difference_cols hinzu, sodass ausgewählte Spalten ohne Primärschlüssel die Erkennung von Aktualisierungen steuern, und bestätige das gewählte Verhalten für den Fall, dass sich keine der angeforderten Spalten mit den verfügbaren Spalten überschneidet.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Feature Request / Improvement
Currently, when detecting which rows should be updated in the upsert, all non-primary key columns are iterated over, converted to a Python type, and compared. This can be time-consuming and memory-consuming (e.g., with complex columns containing JSON data in struct, list, ...).
If the user knows that a change has occurred to specific columns, it would be a huge performance improvement to just iterate over these columns. Or if the user has a column that implies that any change has occurred (e.g., a hash of the data)
For example:
Hash column: if the user has the possibility to create a column containing a hash for each row, then upsert has to look only at this column when detecting changes. This way, pyiceberg doesn't have to convert all columns to Python type and compare them.
Proposition:
Update the function get_rows_to_update to also accept an optional parameter difference_cols. Update the upsert methods with this parameter and pass it to the get_rows_to_update.
In case there is no intersection between non-primary key columns and difference_cols, pyiceberg can either raise and error or it can fall-back to the default behaviour (iterating over all non-PK columns).
Usage:
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField, StringType
import pyarrow as pa
schema = Schema(
NestedField(1, "city", StringType(), required=True),
NestedField(2, "inhabitants", IntegerType(), required=True),
# Mark City as the identifier field, also known as the primary-key
identifier_field_ids=[1]
)
tbl = catalog.create_table("default.cities", schema=schema)
arrow_schema = pa.schema(
[
pa.field("city", pa.string(), nullable=False),
pa.field("inhabitants", pa.int32(), nullable=False),
]
)
# Write some data
df = pa.Table.from_pylist(
[
{"city": "Amsterdam", "inhabitants": 921402},
{"city": "San Francisco", "inhabitants": 808988},
{"city": "Drachten", "inhabitants": 45019},
{"city": "Paris", "inhabitants": 2103000},
],
schema=arrow_schema
)
tbl.append(df)
df = pa.Table.from_pylist(
[
# Will be updated, the inhabitants has been updated
{"city": "Drachten", "inhabitants": 45505},
# New row, will be inserted
{"city": "Berlin", "inhabitants": 3432000},
# Ignored, already exists in the table
{"city": "Paris", "inhabitants": 2103000},
],
schema=arrow_schema
)
upd = tbl.upsert(df, difference_cols=["inhabitants"])
I have already prepared how it can look in my fork 59d18337a61b106566c2e6a432b7d4899ca7f334.
If the proposition is accepted, I can prepare the PR.
- Vorherrschende Sprache
- Python
- Sterne
- 1.1k
- Forks
- 606
- Ø Merge
- 1 T. 11 Std.
- Gemergte PRs (30 T.)
- 76
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Kein Beitragsleitfaden
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/iceberg-python
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseEvtl. vergeben @Soumo-git-hub hat das heute übernommen. Offenkind:bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
apache/iceberg-python#4073 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
apache/iceberg-python#4010 · 3 Kommentare · 1 Reaktion ·
Maintainer antworten meist innerhalb von 1 Tag
-
to_bytes silently rescales a Decimal with a negative scaleEvtl. vergeben @Rodrigo-Palma hat das vor 17 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/iceberg-python#3996 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationEvtl. vergeben @ghoshp83 hat das vor 17 Tagen übernommen. Offenbug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
apache/iceberg-python#3979 ·
Maintainer antworten meist innerhalb von 1 Tag
-
FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystemEvtl. vergeben @krishnakaanchan-png hat das vor 34 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/iceberg-python#3885 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in apache/iceberg-python
Ähnliche Issues
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Evtl. vergeben @mayureshsonawane17 hat das heute übernommen. OffenWaiting for: Product Owner
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
Maintainer antworten meist innerhalb von 5 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 88/100
bojieli/ai-agent-book#1169 ·
Maintainer antworten meist innerhalb von 1 Tag
-
priority:low ready-for-dev
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
OpenHands/extensions#738 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
micronaut-projects/micronaut-core#13677 ·
Maintainer antworten meist innerhalb von 1 Tag