Autoloop PR/issue should report cumulative performance improvement
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Problem
Autoloop programs that optimize a metric over many iterations don't report the cumulative improvement anywhere visible — not in the PR body, not in the issue, not in the state file summary. A reader of tsessebe PR #297 or issue #189 can see the current best metric but has no idea where the program started or how much progress has been made.
This was observed on the tsb-perf-evolve program (41 iterations, fitness ratio optimization). The PR body says "Current best metric: 21.048" but never mentions the starting value or total improvement.
Root causes
Three gaps in autoloop.md:
1. No Initial Metric in Machine State
The state file schema tracks Best Metric but not the baseline metric from the first accepted iteration. Without this, there's no reference point to compute cumulative improvement.
Fix: Add an Initial Metric field to the Machine State table schema. Set it on the first accepted iteration and never overwrite it.
2. PR body template doesn't include cumulative improvement
Step 5c says to update the PR body with "the latest metric and a summary of the most recent accepted iteration" but never specifies showing start-to-finish improvement.
Fix: Update the Step 5c PR body template to include: "Fitness: {best_metric} (started at {initial_metric}, {improvement_pct}% improvement)" or similar.
3. pending-ci fitness values never flow back into PR/issue
When the sandbox can't run the evaluation command (e.g., bun not available), iterations are pushed as pending-ci. After CI runs the benchmark and produces a fitness number, that result never flows back into the PR body, issue status comment, or state file. The PR stays stuck showing the last sandbox-measured metric.
Fix: Either (a) add a post-CI callback step that updates the PR body/state file with the CI-measured fitness, or (b) at minimum, have the next iteration read CI results from the previous run and update the state file retroactively before proposing a new change.
Expected behavior
A reader of any Autoloop PR or program issue should be able to see at a glance:
- Where the metric started (initial/baseline value)
- Where it is now (current best)
- Total improvement (absolute delta and percentage)
- A brief improvement trajectory (e.g., in the PR body or status comment)
References
- Observed on: githubnext/tsessebe PR #297, issue #189
- State file:
tsb-perf-evolve.md - Workflow definition:
.github/workflows/autoloop.md
- Ngôn ngữ chính
- Python
- Star
- 74
- Fork
- 6
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của githubnext/autoloop
-
Autoloop template should exclude language-specific dev files from protected-files by defaultĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
githubnext/autoloop#71 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
githubnext/autoloop#54 ·
-
TESTĐang mở
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 15/100
githubnext/autoloop#76 · 1 bình luận ·
-
ReleasesĐang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
githubnext/autoloop#72 · 1 bình luận ·
-
Install instructions should remind user to set COPILOT_GITHUB_TOKENCó thể đã có người làm @mrjf đã nhận 147 ngày trước. Đang mở
githubnext/autoloop#65 · 1 reaction · 2 người được giao ·
Tất cả issue của githubnext/autoloop
Issue tương tự
-
bug needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
debpalash/VoiceStudio#2624 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Make Catch2 optional when `RDK_BUILD_CPP_TESTS=OFF`Có thể đã có người làm @pechersky đã nhận hôm nay. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 2 ngày
-
There are a few redundant calls to `fdesc._setCloseOnExec()`Có thể đã có người làm @gudnimg đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày