Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Fully lazy imperative function evaluation

オープン
#499 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
python
領域
data

調査の方向性

リンクされている axis=None の cumsum/cumprod 実装を読み、lazyexpr の計算機構、LazyUDFs、matmul のチャンク処理を追跡してください。fancy-indexing のコンテキストについては issue #441 を確認してください。完了条件は、関数がネストした matmul/sum 式を含め、要求された出力チャンクを遅延評価し、完全な結果の一時データを作成せずに、numexpr の優先を維持し、削減処理の繰り返しを避けることです。

索引モデルが issue の本文から書いたものです。

説明

enhancement sustain-2026

In order to make function evaluation properly lazy, all functions (e.g. matmul) should be implemented as LazyUDFs.

Secondly, the lazyexpr machinery of compute should loop over chunks of the result. Each function must then decide what slices of the operands are necessary to form the corresponding result chunk (as matmul currently does internally). Then, upon evaluation, although the expression evaluates term-by-term, it does not compute the full result for e..g matmul before proceeding to the next term, but only for the necessary chunk(s) of output. Thus there is a higher chance of cache hits (not the case currently for eager execution of linalg and reductions).
Something in this spirit has been implemented for cumsum/cumprod when axis=None (see this code ):

# Special case for cumulative operations with axis = None
        if reduce_args["axis"] is None and reduce_op in {ReduceOp.CUMULATIVE_PROD, ReduceOp.CUMULATIVE_SUM}:
            # res_out_ is just None, out set to all 0s (sum) or 1s (prod)
            out, res_out_ = convert_none_out(dtype, reduce_op, reduced_shape)
            # reduced_shape is just one-element tuple
            chunklen = out.chunks[0] if hasattr(out, "chunks") else chunks[-1]
            carry = 0
            for cidx in range(0, reduced_shape[0] // chunklen):
                slice_starts = np.unravel_index(cidx * chunklen, shape)
                slice_stops = np.unravel_index((cidx + 1) * chunklen, shape)
                cslice = tuple(
                    slice(start, stop) for start, stop in zip(slice_starts, slice_stops, strict=True)
                )
                _get_chunk_operands(operands, cslice, chunk_operands, shape)
                result, _ = _get_result(expression, chunk_operands, ne_args, where)
                result = np.require(result, requirements="C")
                if reduce_op == ReduceOp.CUMULATIVE_SUM:
                    res = np.cumulative_sum(result, axis=None) + carry
                else:
                    res = np.cumulative_prod(result, axis=None) * carry
                carry = res[-1]
                out[cidx * chunklen + include_initial : (cidx + 1) * chunklen + include_initial] = res

It should also be implemented for fancy-indexing with slice (see https://github.com/Blosc/python-blosc2/issues/441).

Should be possible to handle even something like "matmul(sum(a,axis=1), b)" by passing the desired slice from matmul->sum, which treats the asked-for-slice as a desired output slice and handles accordingly.

This would avoid large in-memory temporaries for example when calculating from eagerly executed linear algebra functions e.g. in "matmul(a, b) + b".

Problems:
1 - reductions could be a problem since for example "sum(a) + a" for the chunks of output would recalculate the scalar "sum(a)" for each chunk of output. This could be avoided by using some kind of result cache for reductions.
2 - naturally, numexpr would always be faster since it compiles the expression into bytecode. Thus we should make sure that, when possible, we still use numexpr preferentially (for most elementwise funcs).

主要言語
Python
スター
211
フォーク
66
平均マージ
2日 1時間
マージ済み PR(30日)
15

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Blosc/python-blosc2 のほかの issue

Blosc/python-blosc2 の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。