[Request] Release optimized skill artifacts for additional methods / models / harnesses
Maintainers usually reply within 4 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
Research direction
The currently released artifacts are in ckpt/ as compact best_skill.md files; start by reviewing those files and the Table 1 cells listed in the issue. Done means publishing the requested additional artifacts and, where available, their accepted-edit histories and selection/test scores.
Written by the indexing model from the issue text.
Description
Hi SkillOpt team,
Thank you for open-sourcing the best GPT-5.5 SkillOpt checkpoints under ckpt/
(https://github.com/microsoft/SkillOpt/tree/main/ckpt) — they have been very
helpful for getting started.
Context
I'm doing an analysis of the learned skills and would like to build on SkillOpt.
Since optimizing these skills is token-expensive and I'm compute/budget-constrained,
re-running the full loop to reproduce every cell in Table 1 is impractical on my end.
Because the exported artifacts are compact best_skill.md files (per the paper, mostly
<2,000 tokens), releasing them should be low-cost on your side and would let the
community study and compare the skills directly, and improve reproducibility.
Request
Would it be possible to release the exported skill artifacts (the final best_skill.md,
and ideally the accepted-edit history / selection & test scores if available) for the
following cells?
1. Direct-chat artifacts — one skill per benchmark
(SearchQA, SpreadsheetBench, OfficeQA, DocVQA, LiveMath, ALFWorld):
| Method | GPT-5.5 | GPT-5.4 | GPT-5.4-mini | GPT-5.4-nano | GPT-5.2 | Qwen3.5-4B | Qwen3.6-35B-A3B |
|---|---|---|---|---|---|---|---|
| SkillOpt | ✅ done | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ |
| Trace2Skill | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ |
| GEPA | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ | ⬜ |
2. Harness artifacts (GPT-5.5) — 5 benchmarks, no ALFWorld:
| Method | Codex harness | Claude Code harness |
|---|---|---|
| SkillOpt | ⬜ | ⬜ |
| EvoSkill | ⬜ | ⬜ |
Note: I limited the EvoSkill request to the Codex / Claude Code harnesses because,
if I read Table 1 correctly, EvoSkill is only reported there and not in the
direct-chat rows. Please correct me if direct-chat EvoSkill artifacts also exist.
Priority (in case a full release is too much at once)
If releasing everything is impractical, an incremental release in this order would
already be very useful to me:
- SkillOpt across the remaining 6 direct-chat models (the headline artifacts).
- GEPA and Trace2Skill across all 7 direct-chat models (for cross-method comparison).
- The Codex / Claude Code harness artifacts (SkillOpt + EvoSkill).
Happy to help however useful — e.g. verifying/organizing the released files, or writing
a short README for the extra checkpoints — and I will of course cite the paper.
Thanks again for the great work and for open-sourcing it!
- Dominant language
- Python
- Stars
- 17.3k
- Forks
- 1.6k
- Avg merge
- 8d 16h
- Merged PRs (30d)
- 11
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/SkillOpt
-
项目还在迭代嘛?Open
Difficulty 5/5 Over a week Newbie friendliness 10/100
microsoft/SkillOpt#291 · 1 comment ·
Maintainers usually reply within 4 days
-
Difficulty 5/5 Over a week Newbie friendliness 42/100
microsoft/SkillOpt#288 · 1 comment ·
Maintainers usually reply within 4 days
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
microsoft/SkillOpt#286 · 3 comments ·
Maintainers usually reply within 4 days
-
Difficulty 5/5 Over a week Newbie friendliness 28/100
microsoft/SkillOpt#283 · 1 comment ·
Maintainers usually reply within 4 days
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
Maintainers usually reply within 4 days
All issues in microsoft/SkillOpt
Similar issues
-
deployment release-lag
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
nolte/kamerplanter#2047 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NousResearch/hermes-plugin-claude-subscription-directsdk#94 ·
Maintainers usually reply within 1 day
-
namespace operations
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
EclipseFdn/open-vsx.org#13702 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day