Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Tests] Add CI integration for LongMemEval benchmark

Open
#4 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
55/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Quiet
Tech stack
cpp, github-actions
Domain
ci-cd, testing

Research direction

Start with tools/longmemeval-mini/benchmark.py, tools/longmemeval-mini/llama_cli_ab_test.py, and semantic_memory_accuracy_smoke.py to understand the existing benchmark entry points and thresholds. Create .github/workflows/longmemeval.yml with the specified PR and manual triggers, model caching, test runs, and PR results reporting; document it in tests/README.md and verify the workflow blocks or reports failures as intended.

Written by the indexing model from the issue text.

Description

stale

Summary

Integrate the LongMemEval benchmark into CI/CD for regression detection.

Current State

  • benchmark.py - Rule-based proxy benchmark
  • llama_cli_ab_test.py - Real LLM A/B testing
  • semantic_memory_accuracy_smoke.py - Accuracy-focused tests
  • Not currently running in CI

Implementation Plan

  1. Create GitHub Actions workflow .github/workflows/longmemeval.yml
  2. Select lightweight model for CI (e.g., LFM2.5-1.2B Q4_K_M)
  3. Set success thresholds:
    • Memory ON accuracy > Memory OFF accuracy
    • No regressions in any category
  4. Add caching for downloaded models

Workflow Design

Trigger: [pull_request, workflow_dispatch]
Steps:
  1. Build llama-cli with semantic memory
  2. Download/cache test model
  3. Run semantic_memory_smoke.py
  4. Run semantic_memory_accuracy_smoke.py --fail-on-no-lift
  5. Post results as PR comment

Acceptance Criteria

  • Workflow runs on every PR
  • Results posted as PR comment
  • Failed tests block merge (optional)
  • Documentation in tests/README.md

Related Files

  • tools/longmemeval-mini/benchmark.py
  • tools/longmemeval-mini/llama_cli_ab_test.py
Dominant language
C++
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from jose-compu/funes.cpp

All issues in jose-compu/funes.cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.