A new ExpressionRuntime (Lark grammar compile) per evaluated expression makes evaluation ~5x slower
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 86/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- performance, testing
Research direction
Start with src/data_factory_testing_framework/models/_data_factory_element.py and inspect how ExpressionRuntime is created, then read ExpressionRuntime.init and the grammar compilation in expression_transformer.py. Reproduce the timing comparison from the issue and run the project’s pipeline tests. Done means evaluations reuse runtime setup without per-expression grammar compilation and retain their existing behavior.
Written by the indexing model from the issue text.
Description
Every DataFactoryElement.evaluate creates a new ExpressionRuntime, and every ExpressionRuntime compiles the Lark grammar of the expression language again. Evaluating one expression costs about 5x what the evaluation itself takes.
Environment: data-factory-testing-framework 1.4.2 (also present on main at 11b9f27), Python 3.12.13, .NET 10.0.12, Linux.
Repro
import time
from data_factory_testing_framework._expression_runtime.expression_runtime import ExpressionRuntime
from data_factory_testing_framework.models import DataFactoryElement
from data_factory_testing_framework.state import PipelineRunState
state = PipelineRunState([], [])
element = DataFactoryElement("@concat('a', 'b')")
start = time.perf_counter()
for _ in range(50):
element.evaluate(state)
per_call = (time.perf_counter() - start) / 50
runtime = ExpressionRuntime()
start = time.perf_counter()
for _ in range(50):
runtime.evaluate("@concat('a', 'b')", state)
shared = (time.perf_counter() - start) / 50
start = time.perf_counter()
ExpressionRuntime()
construct = time.perf_counter() - start
print(f"DataFactoryElement.evaluate: {per_call * 1000:.1f} ms per expression")
print(f"shared ExpressionRuntime: {shared * 1000:.1f} ms per expression")
print(f"ExpressionRuntime(): {construct * 1000:.1f} ms")
Output
DataFactoryElement.evaluate: 48.0 ms per expression
shared ExpressionRuntime: 9.8 ms per expression
ExpressionRuntime(): 27.9 ms
On a real suite (about 30 pipeline tests, each run evaluating dozens of expressions) this was 109 s; sharing one runtime brought it to 11 s.
Cause
_data_factory_element.py#L29 builds a runtime per call:
expression_runtime = ExpressionRuntime()
and ExpressionRuntime.__init__ builds an ExpressionTransformer, which compiles the grammar (expression_transformer.py#L108).
Suggested fix: create the runtime once (module-level or a cached factory) and reuse it; it holds no per-evaluation state. As a workaround we patch data_factory_testing_framework.models._data_factory_element.ExpressionRuntime to return a shared instance.
- Dominant language
- Python
- Stars
- 135
- Forks
- 43
- Avg merge
- 1h 14m
- Merged PRs (30d)
- 1
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/data-factory-testing-framework
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
IfCondition evaluates the expressions of the branch not taken (case-sensitive "activities" check)Open
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
DataFactoryTestingFrameworkExpressionsEvaluator adds status property even if None is set.Possibly taken @LeonardHd claimed this 434 days ago. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
All issues in microsoft/data-factory-testing-framework
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Juniper/ansible-junos-stdlib#904 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
pollen-robotics/reachy_mini#1457 ·
Maintainers usually reply within 1 day
-
area:runtime good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
WATonomous/wato_f1tenth#39 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
FireDynamics/fdsreader#123 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100