Mining is limited to programmatic success checks — intent-level improvements get dropped or Goodharted into shallow proxy rules
Los mantenedores suelen responder en 6 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 30/100
Línea de trabajo
Empieza por rastrear las comprobaciones de éxito basadas en regex del task miner y el flujo de mining/replay descrito en la issue. Compara los enfoques propuestos divergence-point, clean-context replay y baseline-instance, y después revisa la respuesta de los maintainers para determinar si se acepta un estándar de evidencia basado en judge; la issue solo estará completa cuando se acuerden un diseño concreto y los criterios de aceptación.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
ORIGIN & MOTIVATION
I'm exploring using SkillOpt as a step toward "on-the-job training" for AI Employees learning a specific job role from real work sessions. This issue is about what looks like a structural ceiling on that goal in the mining stage. (Related but distinct from #67, which covers reward-hacking on the gate side; this is about the miner side.)
Problem
The task miner only retains a candidate task if success can be expressed as a programmatic check (regex-style assertions). That means behaviors like "understand my intent better" don't compress into a check, so they are either:
- dropped the most valuable improvement signal never enters the pipeline, or
- mangled into a shallow proxy a phrasing/format preference observed in one session gets distilled into a hyper-literal rule that is then robotically applied to all future outputs. Goodhart's law in miniature: the check becomes the target.
I've observed the second failure mode in my own runs: a one-off summary phrasing preference became a rigid rule stamped onto every future summary.
If mining is restricted to regex-expressible outcomes, the system can only ever improve the regex-expressible slice of agent behavior, which is a small subset of what users actually repeat-and-rephrase about in real sessions.
Possible directions (heuristics, not designs)
- Divergence-point detection: mine topic-shift structure in transcripts (user descends into a rabbit hole on a sub-issue → resolves it → conversation returns to the main thread) as natural task boundaries and implicit failure signals, instead of requiring a programmatic outcome check.
- Clean-context replay comparison: a second instance with fresh context attempts the reconstructed task; a judge compares its output against what the user ultimately accepted in the original transcript, rather than against a regex proxy.
- Baseline-instance divergence: an instance seeded only with the user's global rules/skills replays the session's user prompts in order until its output clearly and definitely diverges from the transcript — the divergence point marks where a learnable behavior lives, and bounds the task to mine.
Question for maintainers
Is there interest in supporting judge-based (non-programmatic) success checks in mining/replay, and what evidence standard would you consider acceptable for gating on them?
- Lenguaje dominante
- Python
- Estrellas
- 18k
- Forks
- 1.7k
- Merge medio
- 6 d 18 h
- PR fusionados (30 d)
- 12
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/SkillOpt
-
Same unconfined `predictions/<item_id>` path pattern remains in four benchmark rolloutsPosiblemente ocupada @adongwanai la tomó hace 1 día. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
Los mantenedores suelen responder en 6 días
-
项目还在迭代嘛?Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 10/100
microsoft/SkillOpt#291 · 1 comentario ·
Los mantenedores suelen responder en 6 días
-
Release cut for the Aug adopt/webui hardening + 2 residual staging gapsPosiblemente ocupada @RohithPariki la tomó hace 18 días. Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 42/100
microsoft/SkillOpt#288 · 2 comentarios ·
Los mantenedores suelen responder en 6 días
-
skillopt-sleep Codex harvest ingests its own headless replay sessionsPosiblemente ocupada @kaluli123123 la tomó hace 19 días. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
microsoft/SkillOpt#286 · 3 comentarios ·
Los mantenedores suelen responder en 6 días
-
Proposal: optional offline typed-decision judge for rubric scoring (SemIf / NanoJev pattern)Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 28/100
microsoft/SkillOpt#283 · 1 comentario ·
Los mantenedores suelen responder en 6 días
Todos los issues de microsoft/SkillOpt
Issues similares
-
HTML: <template> content is extracted as document textPosiblemente ocupada @ryanmeowy la tomó hoy. Abiertobug html
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
docling-project/docling#4714 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
[BUG] Qdrant RAG client applies score_threshold to raw cosine similarity, not the 0-1 score it returnsPosiblemente ocupada @roydonsequeira la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
ashhart/TensorFold#536 ·
Los mantenedores suelen responder en 1 día
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 62/100
NousResearch/hermes-agent#135440 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Deepak3699/Ai_Mentor#244 ·
Los mantenedores suelen responder en 1 día