[Bug] Multi-trace files collapse into one eval case
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 78/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- testing-qa
Research direction
Start at BaseEvaluator._build_eval_set_from_tracing_json() and read the focused regression coverage in tests/test_evaluator.py. Reproduce the multi-trace, out-of-order-span failure first, then verify that each trace becomes an isolated EvalCase with ordered spans and that the EvalSet timestamp uses the earliest generated case timestamp. Run the evaluator tests, Ruff, and git diff --check when done.
Written by the indexing model from the issue text.
Description
Bug description
BaseEvaluator._build_eval_set_from_tracing_json() groups spans by trace_id, but currently keeps conversation, metadata, and timestamps outside the per-trace loop and appends only one EvalCase after the loop.
For a tracing file containing multiple traces, this can collapse all traces into one case, attach metadata from the last trace, and produce an empty conversation when spans are not already ordered by start_time.
Minimal reproduction
On current main (43f958d), a tracing file with two traces and out-of-order call_llm spans produces:
len(eval_set.eval_cases) == 1instead of2- the only case uses the second trace's
app_nameanduser_id conversation == []
The focused regression fails deterministically with:
assert len(eval_set.eval_cases) == 2
E assert 1 == 2
Expected behavior
Could you confirm whether the intended conversion semantics are:
- each
trace_idbecomes one isolatedEvalCase; - spans are ordered by
start_timewithin that trace; - conversation, tool calls, and session metadata never cross trace boundaries;
- the
EvalSettimestamp is the earliest generated case timestamp?
Validated local fix
A local one-commit patch implements the behavior above and currently passes:
tests/test_evaluator.py— 3 passed- Ruff 0.11.12 check and format
git diff --check
I have not opened a PR yet because this changes the public mapping between tracing files and evaluation cases. If the semantics above are intended, I can submit the tested patch.
AI assistance
The investigation and candidate patch were developed with AI assistance. I verified the failure on a clean origin/main worktree with an isolated Python bytecode cache and reviewed the trace-boundary semantics.
- Dominant language
- Python
- Stars
- 345
- Forks
- 99
- Avg merge
- 11h 38m
- Merged PRs (30d)
- 89
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from volcengine/veadk-python
-
Difficulty 2/5 1-3 hours Newbie friendliness 92/100
volcengine/veadk-python#1154 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
volcengine/veadk-python#901 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
volcengine/veadk-python#549 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 55/100
volcengine/veadk-python#531 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 58/100
volcengine/veadk-python#508 · 2 comments ·
Maintainers usually reply within 1 day
All issues in volcengine/veadk-python
Similar issues
-
upstream update
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
conan-io/conan-center-index#31098 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
john-kurkowski/tldextract#382 ·
-
comp/tools duplicate P2 sweeper:risk-compatibility tool/mcp type/bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
NousResearch/hermes-agent#132042 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#13092 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 85/100
feder-cr/invisible_playwright_mcp#1408 ·
Maintainers usually reply within 1 day