[Tests] Add CI integration for LongMemEval benchmark
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
Research direction
Start with tools/longmemeval-mini/benchmark.py, tools/longmemeval-mini/llama_cli_ab_test.py, and semantic_memory_accuracy_smoke.py to understand the existing benchmark entry points and thresholds. Create .github/workflows/longmemeval.yml with the specified PR and manual triggers, model caching, test runs, and PR results reporting; document it in tests/README.md and verify the workflow blocks or reports failures as intended.
Written by the indexing model from the issue text.
Description
Summary
Integrate the LongMemEval benchmark into CI/CD for regression detection.
Current State
benchmark.py- Rule-based proxy benchmarkllama_cli_ab_test.py- Real LLM A/B testingsemantic_memory_accuracy_smoke.py- Accuracy-focused tests- Not currently running in CI
Implementation Plan
- Create GitHub Actions workflow
.github/workflows/longmemeval.yml - Select lightweight model for CI (e.g., LFM2.5-1.2B Q4_K_M)
- Set success thresholds:
- Memory ON accuracy > Memory OFF accuracy
- No regressions in any category
- Add caching for downloaded models
Workflow Design
Trigger: [pull_request, workflow_dispatch]
Steps:
1. Build llama-cli with semantic memory
2. Download/cache test model
3. Run semantic_memory_smoke.py
4. Run semantic_memory_accuracy_smoke.py --fail-on-no-lift
5. Post results as PR comment
Acceptance Criteria
- Workflow runs on every PR
- Results posted as PR comment
- Failed tests block merge (optional)
- Documentation in
tests/README.md
Related Files
tools/longmemeval-mini/benchmark.pytools/longmemeval-mini/llama_cli_ab_test.py
- Dominant language
- C++
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from jose-compu/funes.cpp
-
stale
Difficulty 5/5 Over a week Newbie friendliness 25/100
jose-compu/funes.cpp#12 ·
-
stale
Difficulty 4/5 3-5 days Newbie friendliness 52/100
jose-compu/funes.cpp#11 ·
-
stale
Difficulty 4/5 3-5 days Newbie friendliness 48/100
jose-compu/funes.cpp#10 ·
-
stale
Difficulty 4/5 3-5 days Newbie friendliness 55/100
jose-compu/funes.cpp#9 ·
-
stale
Difficulty 4/5 3-5 days Newbie friendliness 45/100
jose-compu/funes.cpp#8 ·
All issues in jose-compu/funes.cpp
Similar issues
-
area/actorsystem bug tsan
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
ydb-platform/ydb#54282 ·
Maintainers usually reply within 1 day
-
bug needs triage
Difficulty 1/5 Under an hour Newbie friendliness 90/100
project-chip/connectedhomeip#74434 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
tenstorrent/tt-metal#58057 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
maplibre/maplibre-native#4690 ·
Maintainers usually reply within 1 day
-
comp-query-execution
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
ClickHouse/ClickHouse#122569 ·
Maintainers usually reply within 1 day