Evaluation traces
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start by reviewing the evaluation setup in the repository and the paper's Table 1 results. Determine whether the evaluation traces or disaggregated scores are available to package and publish; done means one of those artifacts is publicly shared with enough context to reproduce the table.
Written by the indexing model from the issue text.
Description
Hi,
This is a really interesting dataset! Would you be able to share the evaluation traces used to compile the results in Table 1? Or the disaggregated scores from the table.
Thanks!
- Dominant language
- Python
- Stars
- 122
- Forks
- 11
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/delegate52
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
microsoft/delegate52#6 ·
-
Image domain Open
Difficulty 1/5 Under an hour Newbie friendliness 35/100
microsoft/delegate52#5 · 1 comment ·
All issues in microsoft/delegate52
Similar issues
-
triage/confirmed
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
agentscope-ai/agentscope#2775 ·
-
comp/desktop P3 type/bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
NousResearch/hermes-agent#118866 ·
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 90/100
apache/cloudstack#14222 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100