flight-sql-jdbc-driver: add option to rely on Arrow Flight SQL Bulk Ingestion for batched inserts
Maintainer antworten meist innerhalb von 3 Tagen
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 35/100
Rechercherichtung
Beginnen Sie mit der Überprüfung des Update-Pfads für Prepared Statements des JDBC-Treibers und der Flight SQL Bulk Ingestion RPCs, insbesondere DoPut(CommandStatementIngest) und ActionCreatePreparedStatementRequest. Definieren Sie, wie eine Option geeignete, ausschließlich einfügenden Batch-Statements identifiziert und sie über Bulk Ingestion weiterleitet, bei gleichzeitiger Kompatibilität mit anderen Prepared Statements.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Describe the enhancement requested
Arrow Flight SQL has a feature for ingesting massive datasets, Bulk Ingestion: https://github.com/apache/arrow/issues/38255
It would be beneficial to use those special RPC methods for batched prepared statement calls when the prepared statement is strictly for inserting data.
E.g. when Spark is used for writing data, it generates a simple SQL query like "INSERT INTO table(field1, field2, ...) VALUES (?, ?, ...)", creates a prepared statement, and then uses the prepared statement update RPC method to insert the rows of the dataset. If this feature is implemented, it would be possible for the driver to instead use the DoPut(CommandStatementIngest) command associated with the Bulk Ingestion feature instead.
There are Arrow Flight SQL server implementations that work like this: when a DoAction(ActionCreatePreparedStatementRequest) is executed, the server creates up to two versions of the data structure underlying the instance of the PreparedStatement. One is a handle to a full-scale query engine execution procedure (e.g. DataFusion's logical plan), and another is a handle to a very simple procedure that just stores the received record batches in the storage - of course, the second procedure is only possible to be generated when the query is of a certain form; like the one used by Spark. The point is that this simple procedure also works without overheads associated with the more general interface of prepared statement API - for example, it does not need to do a transposition of PreparedStatement parameters into record batches.
I think that it should be possible to move this logic for deciding to use Bulk Ingestion into the jdbc driver.
Usecase for this integration is this: developers of Arrow Flight SQL servers could implement bulk ingestion command handlers and avoid implementing special logic for handling batched inserts in a special manner. Then the client would use this newly introduced driver option to allow the driver to decide to use the bulk ingestion RPC methods for inserting data.
- Vorherrschende Sprache
- Java
- Sterne
- 97
- Forks
- 158
- Ø Merge
- 9 T. 4 Std.
- Gemergte PRs (30 T.)
- 13
Entwicklungsumgebung
- Enthält ein Dockerfile oder eine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/arrow-java
-
DictionaryEncoder.decode accepts out-of-range dictionary indicesEvtl. vergeben @Arawoof06 hat das vor 46 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 74/100
apache/arrow-java#1261 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
ArrowFlightJdbcArray.getArray(index, count) can read past the end of the array sliceEvtl. vergeben @Arawoof06 hat das vor 75 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/arrow-java#1236 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
[Java][JDBC] ClobConsumer writes past VarCharVector data buffer for large CLOBsEvtl. vergeben @Arawoof06 hat das vor 80 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/arrow-java#1230 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
Byte-array elements leak in `FromSchemaByteArray()`Evtl. vergeben @PG1204 hat das vor 71 Tagen übernommen. OffenType: bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
apache/arrow-java#1205 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
[Java][IPC] AbstractCompressionCodec.compress() writes prefix=0 for empty buffers, incompatible with C++/Python readersEvtl. vergeben @PG1204 hat das vor 71 Tagen übernommen. OffenType: bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
apache/arrow-java#1196 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 3 Tagen
Alle Issues in apache/arrow-java
Ähnliche Issues
-
test(setup): GitHub configuration tests fail when the temp path is long enough for YAML foldingOffenbug good first issue help wanted priority medium size S
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
martin-francois/symphony-trello#776 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Console.printHexOffengood first issue kernel
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
JackFurton/who-would-build-a-kernel-in-java#33 · 2 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
objectionary/eo#9182 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
objectionary/hone-maven-plugin#1293 ·
Maintainer antworten meist innerhalb von 1 Tag