Autoloop PR/issue should report cumulative performance improvement
Assessment
This issue has not been assessed yet.
Description
Problem
Autoloop programs that optimize a metric over many iterations don't report the cumulative improvement anywhere visible — not in the PR body, not in the issue, not in the state file summary. A reader of tsessebe PR #297 or issue #189 can see the current best metric but has no idea where the program started or how much progress has been made.
This was observed on the tsb-perf-evolve program (41 iterations, fitness ratio optimization). The PR body says "Current best metric: 21.048" but never mentions the starting value or total improvement.
Root causes
Three gaps in autoloop.md:
1. No Initial Metric in Machine State
The state file schema tracks Best Metric but not the baseline metric from the first accepted iteration. Without this, there's no reference point to compute cumulative improvement.
Fix: Add an Initial Metric field to the Machine State table schema. Set it on the first accepted iteration and never overwrite it.
2. PR body template doesn't include cumulative improvement
Step 5c says to update the PR body with "the latest metric and a summary of the most recent accepted iteration" but never specifies showing start-to-finish improvement.
Fix: Update the Step 5c PR body template to include: "Fitness: {best_metric} (started at {initial_metric}, {improvement_pct}% improvement)" or similar.
3. pending-ci fitness values never flow back into PR/issue
When the sandbox can't run the evaluation command (e.g., bun not available), iterations are pushed as pending-ci. After CI runs the benchmark and produces a fitness number, that result never flows back into the PR body, issue status comment, or state file. The PR stays stuck showing the last sandbox-measured metric.
Fix: Either (a) add a post-CI callback step that updates the PR body/state file with the CI-measured fitness, or (b) at minimum, have the next iteration read CI results from the previous run and update the state file retroactively before proposing a new change.
Expected behavior
A reader of any Autoloop PR or program issue should be able to see at a glance:
- Where the metric started (initial/baseline value)
- Where it is now (current best)
- Total improvement (absolute delta and percentage)
- A brief improvement trajectory (e.g., in the PR body or status comment)
References
- Observed on: githubnext/tsessebe PR #297, issue #189
- State file:
tsb-perf-evolve.md - Workflow definition:
.github/workflows/autoloop.md
- Dominant language
- Python
- Stars
- 72
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from githubnext/autoloop
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
githubnext/autoloop#71 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
githubnext/autoloop#54 ·
-
Releases Open
Difficulty 4/5 3-5 days Newbie friendliness 35/100
githubnext/autoloop#72 · 1 comment ·
-
githubnext/autoloop#65 · 1 reaction · 2 assignees ·
-
githubnext/autoloop#63 · 1 reaction · 2 assignees ·
All issues in githubnext/autoloop
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
learningequality/ricecooker#747 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
BSData/horus-heresy-3rd-edition#3171 ·
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
run-llama/llama_index#23199 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
KhronosGroup/glTF-Blender-IO#2769 ·