Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

[Java] Parquet DatasetFileWriter not completely emptying allocator after finishing a write

Offen
#1,255 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 2 Tagen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
35/100
Issue-Typ
Bug
Klarheit
Muss geklärt werden
Aktivitätsstatus
Ruhig
Tech-Stack
java
Bereich
data

Rechercherichtung

Beginne mit dataset/src/test/java/org/apache/arrow/dataset/file/TestDatasetFileWriter.java und dem im Bericht beschriebenen DatasetFileWriter.write-Aufruf. Reproduziere den Parquet-Fall mit einer Zeile mit writerAllocator und seinen untergeordneten Allocatoren und verfolge anschließend, wann der Speicher im Verhältnis zum Schließen von Writer und Allocator den Wert null erreicht. Als erledigt gilt die Untersuchung, wenn festgestellt wurde, ob es sich um eine erwartete Verwendung oder einen Bibliotheksfehler handelt, und die korrekten Hinweise zum Lebenszyklus dokumentiert sind.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Type: usage
Describe the usage question you have. Please include as many useful details as possible.

I'm currently working on a somewhat simple project that processes Parquet files as both input and output:

  • Main thread is a Parquet DatasetFactory that reads the file in batches. Here I have a child readerAllocator.
  • The main thread schedules a fixed amount of worker threads that do the actual processing. They each have their own child workerAllocator.
  • The main thread also schedules a single writer thread that reads (via a custom ArrowReader) a queue of worker thread futures from which to obtain the new VectorSchemaRoot with additional columns. The transfer is done just fine via VectorLoader/Unloader and sent to a child writerArrowReaderAllocator, which is then read by the actual writer's child writerAllocator.

I'll admit that this is my first time using the library and am not completely up to speed with the memory management, so my first reaction to seeing the allocators as an AutoCloseable, was to use Java's try-with-resource syntax to not have to faff about with closing the allocators manually.

Much, much, much more debugging later... I realize this is a terrible idea as the allocator's behavior is not always synced with the different objects (e.g.: DatasetFileWriter) that use them (i.e.: the allocators are not empty by the time the closeable objects are closed).

I would get my output parquet file (first one at least) but then the application would exit abruptly due to closing the allocators too early as they weren't empty.

Insanity set in and I wrote this (absolutely disgusting) piece of code out of desperation (putting it at the very end after the DatasetFileWriter.write call has finished):

boolean childrenEmpty = false;
while (!childrenEmpty) {
    childrenEmpty = true;
    for (BufferAllocator childAllocator : rootAllocator.getChildAllocators()) {
        if (allocator.getAllocatorMemory <= 0) allocator.close();
        else childrenEmpty = false;
    }
}

and... finally... everything started working... no memory errors anymore...

Doing a simple test with a parquet file of 1 row, it would take the writerAllocator roughly anywhere from 50 to 600 iterations until the allocated memory bytes would be down to 0 and then the allocator could be closed without errors.

I'm pretty I've made mistakes in my design given that this is my first time using the library. However, in every example I've seen (whether it be from random google results, the cookbook, or even Claude Opus), the writer's allocator is handled with a try-with-resource. Heck even the class's test closes the allocator without checking if its empty...

So I would just like to know if there is some hidden known side effect that I haven't read about or if there is some usage guidelines I simply haven't followed.

I'm also open to any alternative suggestions on how to handle the closing of the allocators.

Component(s)

OS: RHEL 9.4
Java version: OpenJDK 26.0.1 (compiling to version 25)
Arrow JAR versions: 19.0.0
Default Memory Allocation Manager: Netty
JVM options: --sun-misc-unsafe-memory-access=allow --add-opens=java.base/java.nio=ALL-UNNAMED

Vorherrschende Sprache
Java
Sterne
97
Forks
158
Ø Merge
7 T. 19 Std.
Gemergte PRs (30 T.)
26

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus apache/arrow-java

Alle Issues in apache/arrow-java

Ähnliche Issues

Weitere Issues zu Java

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.