Mining is limited to programmatic success checks — intent-level improvements get dropped or Goodharted into shallow proxy rules
Maintainers usually reply within 4 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
Research direction
Start by tracing the task miner's regex-based success checks and the mining/replay flow described in the issue. Compare the proposed divergence-point, clean-context replay, and baseline-instance approaches, then review the maintainers' response to determine whether a judge-based evidence standard is accepted; the issue is complete only when a concrete design and acceptance criteria are agreed.
Written by the indexing model from the issue text.
Description
ORIGIN & MOTIVATION
I'm exploring using SkillOpt as a step toward "on-the-job training" for AI Employees learning a specific job role from real work sessions. This issue is about what looks like a structural ceiling on that goal in the mining stage. (Related but distinct from #67, which covers reward-hacking on the gate side; this is about the miner side.)
Problem
The task miner only retains a candidate task if success can be expressed as a programmatic check (regex-style assertions). That means behaviors like "understand my intent better" don't compress into a check, so they are either:
- dropped the most valuable improvement signal never enters the pipeline, or
- mangled into a shallow proxy a phrasing/format preference observed in one session gets distilled into a hyper-literal rule that is then robotically applied to all future outputs. Goodhart's law in miniature: the check becomes the target.
I've observed the second failure mode in my own runs: a one-off summary phrasing preference became a rigid rule stamped onto every future summary.
If mining is restricted to regex-expressible outcomes, the system can only ever improve the regex-expressible slice of agent behavior, which is a small subset of what users actually repeat-and-rephrase about in real sessions.
Possible directions (heuristics, not designs)
- Divergence-point detection: mine topic-shift structure in transcripts (user descends into a rabbit hole on a sub-issue → resolves it → conversation returns to the main thread) as natural task boundaries and implicit failure signals, instead of requiring a programmatic outcome check.
- Clean-context replay comparison: a second instance with fresh context attempts the reconstructed task; a judge compares its output against what the user ultimately accepted in the original transcript, rather than against a regex proxy.
- Baseline-instance divergence: an instance seeded only with the user's global rules/skills replays the session's user prompts in order until its output clearly and definitely diverges from the transcript — the divergence point marks where a learnable behavior lives, and bounds the task to mine.
Question for maintainers
Is there interest in supporting judge-based (non-programmatic) success checks in mining/replay, and what evidence standard would you consider acceptable for gating on them?
- Dominant language
- Python
- Stars
- 18k
- Forks
- 1.7k
- Avg merge
- 8d 16h
- Merged PRs (30d)
- 11
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/SkillOpt
-
项目还在迭代嘛?Open
Difficulty 5/5 Over a week Newbie friendliness 10/100
microsoft/SkillOpt#291 · 1 comment ·
Maintainers usually reply within 4 days
-
Release cut for the Aug adopt/webui hardening + 2 residual staging gapsPossibly taken @RohithPariki claimed this 14 days ago. Open
Difficulty 5/5 Over a week Newbie friendliness 42/100
microsoft/SkillOpt#288 · 1 comment ·
Maintainers usually reply within 4 days
-
skillopt-sleep Codex harvest ingests its own headless replay sessionsPossibly taken @kaluli123123 claimed this 15 days ago. Open
Difficulty 3/5 1-2 days Newbie friendliness 72/100
microsoft/SkillOpt#286 · 3 comments ·
Maintainers usually reply within 4 days
-
Difficulty 5/5 Over a week Newbie friendliness 28/100
microsoft/SkillOpt#283 · 1 comment ·
Maintainers usually reply within 4 days
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
Maintainers usually reply within 4 days
All issues in microsoft/SkillOpt
Similar issues
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Possibly taken @mayureshsonawane17 claimed this today. OpenWaiting for: Product Owner
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 5 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
bojieli/ai-agent-book#1169 ·
Maintainers usually reply within 1 day
-
priority:low ready-for-dev
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
OpenHands/extensions#738 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
micronaut-projects/micronaut-core#13677 ·
Maintainers usually reply within 1 day