Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug]: DaskRunner `DaskBagWindowedIterator` materializes entire dataset into memory causing OOM

Aperta
#40,282 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@Gaurav598 ci sta già lavorando.

Dal 10/10/2026.

  • #40283 di @vishalmore90 — chiusa senza merge
  • #40496 di @Gaurav598 — aperta

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
56/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python

Direzione di ricerca

Start in sdks/python/apache_beam/runners/dask/transform_evaluator.py at DaskBagWindowedIterator.iter and review the FIXME around list(self.bag). Investigate the partition-based and delayed-result approaches described in the issue. Done means side-input results are yielded incrementally without materializing the full Dask Bag in client memory, avoiding the reported OOM behavior.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug P2 python
What happened?

In the experimental Python Dask Runner, the DaskBagWindowedIterator (which handles iterators for apache_beam.transforms.sideinputs.SideInputMap) iterates over a Dask Bag by wrapping it in a Python list().

Calling list(self.bag) implicitly triggers a full compute() on the Dask dataset. This blocking operation materializes the entirety of the side input data into the client's local memory. For large side inputs, this completely bypasses Dask's distributed memory management and results in an Out-Of-Memory (OOM) crash, effectively bottlenecking the scalability of pipelines running on Dask.

The code currently includes an explicit FIXME acknowledging this proof-of-concept behavior, but it remains a silent, critical scalability flaw.

Code Pointers / Steps to Reproduce

The issue is located in sdks/python/apache_beam/runners/dask/transform_evaluator.py within the __iter__ method of the DaskBagWindowedIterator class (lines 91-96):

class DaskBagWindowedIterator:
  """Iterator for `apache_beam.transforms.sideinputs.SideInputMap`"""

  bag: db.Bag
  window_fn: WindowFn

  def __iter__(self):
    # FIXME(cisaacstern): list() is likely inefficient, since it presumably
    # materializes the full result before iterating over it. doing this for
    # now as a proof-of-concept. can we can generate results incrementally?
    for result in list(self.bag):
      yield get_windowed_value(result, self.window_fn)

Impact

Any Apache Beam pipeline using DaskRunner that relies on substantial side inputs will crash with OOM errors as soon as the side input data surpasses the available local RAM on the node where the iterator is evaluated. This severely limits the DaskRunner's ability to process real-world distributed datasets and creates a harsh scalability ceiling.

Proposed Solution

The evaluation should generate results incrementally rather than performing a monolithic evaluation. Potential approaches:

  1. Partition-based iteration: Utilize Dask's .map_partitions or .to_delayed() to fetch and yield the underlying data partition-by-partition.
  2. Generators: Instead of eager computation via list(), retrieve delayed results asynchronously and yield them to allow the Python garbage collector to free memory between partition iterations.
Issue Priority

Priority: 2 (default / most bugs should be filed as P2)

Issue Components
  • Component: Python SDK
  • Component: Java SDK
  • Component: Go SDK
  • Component: Typescript SDK
  • Component: IO connector
  • Component: Beam YAML
  • Component: Beam examples
  • Component: Beam playground
  • Component: Beam katas
  • Component: Website
  • Component: Infrastructure
  • Component: Spark Runner
  • Component: Flink Runner
  • Component: Prism Runner
  • Component: Twister2 Runner
  • Component: Hazelcast Jet Runner
  • Component: Google Cloud Dataflow Runner
Lingua principale
Java
Stelle
8.7k
Fork
4.7k
Merge medio
2g 7h
PR unite (30g)
242

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/beam

Tutte le issue di apache/beam

Issue simili

Altre issue su Java

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.