Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

A new ExpressionRuntime (Lark grammar compile) per evaluated expression makes evaluation ~5x slower

Open Beginner friendly
#178 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
86/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start with src/data_factory_testing_framework/models/_data_factory_element.py and inspect how ExpressionRuntime is created, then read ExpressionRuntime.init and the grammar compilation in expression_transformer.py. Reproduce the timing comparison from the issue and run the project’s pipeline tests. Done means evaluations reuse runtime setup without per-expression grammar compilation and retain their existing behavior.

Written by the indexing model from the issue text.

Description

Every DataFactoryElement.evaluate creates a new ExpressionRuntime, and every ExpressionRuntime compiles the Lark grammar of the expression language again. Evaluating one expression costs about 5x what the evaluation itself takes.

Environment: data-factory-testing-framework 1.4.2 (also present on main at 11b9f27), Python 3.12.13, .NET 10.0.12, Linux.

Repro

import time

from data_factory_testing_framework._expression_runtime.expression_runtime import ExpressionRuntime
from data_factory_testing_framework.models import DataFactoryElement
from data_factory_testing_framework.state import PipelineRunState

state = PipelineRunState([], [])
element = DataFactoryElement("@concat('a', 'b')")

start = time.perf_counter()
for _ in range(50):
    element.evaluate(state)
per_call = (time.perf_counter() - start) / 50

runtime = ExpressionRuntime()
start = time.perf_counter()
for _ in range(50):
    runtime.evaluate("@concat('a', 'b')", state)
shared = (time.perf_counter() - start) / 50

start = time.perf_counter()
ExpressionRuntime()
construct = time.perf_counter() - start

print(f"DataFactoryElement.evaluate: {per_call * 1000:.1f} ms per expression")
print(f"shared ExpressionRuntime:    {shared * 1000:.1f} ms per expression")
print(f"ExpressionRuntime():         {construct * 1000:.1f} ms")

Output

DataFactoryElement.evaluate: 48.0 ms per expression
shared ExpressionRuntime:    9.8 ms per expression
ExpressionRuntime():         27.9 ms

On a real suite (about 30 pipeline tests, each run evaluating dozens of expressions) this was 109 s; sharing one runtime brought it to 11 s.

Cause

_data_factory_element.py#L29 builds a runtime per call:

expression_runtime = ExpressionRuntime()

and ExpressionRuntime.__init__ builds an ExpressionTransformer, which compiles the grammar (expression_transformer.py#L108).

Suggested fix: create the runtime once (module-level or a cached factory) and reuse it; it holds no per-evaluation state. As a workaround we patch data_factory_testing_framework.models._data_factory_element.ExpressionRuntime to return a shared instance.

Dominant language
Python
Stars
135
Forks
43
Avg merge
1h 14m
Merged PRs (30d)
1

Getting set up

Open in Codespaces

Starts the project's dev container in your browser, under your own GitHub account.

  • No Dockerfile or Docker Compose file
  • No pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/data-factory-testing-framework

All issues in microsoft/data-factory-testing-framework

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.