Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

A new ExpressionRuntime (Lark grammar compile) per evaluated expression makes evaluation ~5x slower

オープン 初心者向け
#178 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

@florian-burelli がすでに取り組んでいます。

2026年10月7日 から。

  • #182 @florian-burelli による — オープン

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
86/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python

調査の方向性

src/data_factory_testing_framework/models/_data_factory_element.py から始めて ExpressionRuntime がどのように作成されるかを確認し、次に ExpressionRuntime.init と expression_transformer.py での文法のコンパイルを読みます。Issue にある時間比較を再現し、プロジェクトのパイプラインテストを実行します。評価でランタイムのセットアップが再利用され、式ごとに文法をコンパイルせず、既存の動作が維持されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Every DataFactoryElement.evaluate creates a new ExpressionRuntime, and every ExpressionRuntime compiles the Lark grammar of the expression language again. Evaluating one expression costs about 5x what the evaluation itself takes.

Environment: data-factory-testing-framework 1.4.2 (also present on main at 11b9f27), Python 3.12.13, .NET 10.0.12, Linux.

Repro

import time

from data_factory_testing_framework._expression_runtime.expression_runtime import ExpressionRuntime
from data_factory_testing_framework.models import DataFactoryElement
from data_factory_testing_framework.state import PipelineRunState

state = PipelineRunState([], [])
element = DataFactoryElement("@concat('a', 'b')")

start = time.perf_counter()
for _ in range(50):
    element.evaluate(state)
per_call = (time.perf_counter() - start) / 50

runtime = ExpressionRuntime()
start = time.perf_counter()
for _ in range(50):
    runtime.evaluate("@concat('a', 'b')", state)
shared = (time.perf_counter() - start) / 50

start = time.perf_counter()
ExpressionRuntime()
construct = time.perf_counter() - start

print(f"DataFactoryElement.evaluate: {per_call * 1000:.1f} ms per expression")
print(f"shared ExpressionRuntime:    {shared * 1000:.1f} ms per expression")
print(f"ExpressionRuntime():         {construct * 1000:.1f} ms")

Output

DataFactoryElement.evaluate: 48.0 ms per expression
shared ExpressionRuntime:    9.8 ms per expression
ExpressionRuntime():         27.9 ms

On a real suite (about 30 pipeline tests, each run evaluating dozens of expressions) this was 109 s; sharing one runtime brought it to 11 s.

Cause

_data_factory_element.py#L29 builds a runtime per call:

expression_runtime = ExpressionRuntime()

and ExpressionRuntime.__init__ builds an ExpressionTransformer, which compiles the grammar (expression_transformer.py#L108).

Suggested fix: create the runtime once (module-level or a cached factory) and reuse it; it holds no per-evaluation state. As a workaround we patch data_factory_testing_framework.models._data_factory_element.ExpressionRuntime to return a shared instance.

主要言語
Python
スター
135
フォーク
43
平均マージ
1時間 14分
マージ済み PR(30日)
1

環境構築

Codespaces で開く

このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。

  • Dockerfile・Docker Compose ファイルなし
  • プルリクエストのテンプレートなし
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

microsoft/data-factory-testing-framework のほかの issue

microsoft/data-factory-testing-framework の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。