Allow user-defined key/value tags on runs for filtering by test intent
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- typescript
- Domain
- backend-api-design, frontend
Research direction
No files, tests, or entry points are named. Start by tracing run/request creation and the Runs table, then identify the API and persistence paths involved; done means optional tags can be created, edited, returned, displayed, filtered, and preserved across retries or reruns.
Written by the indexing model from the issue text.
Description
Original author: @JaGord
Summary
Please add support for user-defined key/value tags on SCOPE runs and requests so users can filter and group runs by the reason they were created.
Why
When running benchmark matrices, the profile name and task prompt are not always enough to explain the intent of a run. Users need to mark runs with labels such as:
tag: purpose
value: test-agent-kit-branch
or:
tag: criteria_focus
value: vector-index-policy
This would make it easier to answer questions like:
- Which runs were created to test a specific Agent Kit change?
- Which runs were created to validate a specific criterion or failure mode?
- Which runs are part of a rerun batch after a harness fix?
- Which runs should be included or excluded from a comparison table?
Requested capability
Allow users to attach arbitrary key/value tags to a run or request at creation time and edit them later.
Example shape:
{
"tags": {
"purpose": "test-agent-kit-branch",
"criteria_focus": "source_excludes_embedding_from_range_index",
"batch": "rag-chat-rerun-20260626",
"include_in_totals": "true"
}
}
Expected behavior
- Tags are visible in run/request details.
- Tags can be added at request creation time.
- Tags can be edited after creation.
- Runs can be filtered by tag key and key/value pair.
- Tags are returned by the API so automation can group and summarize runs reliably.
- Reports should include tags or link back to tagged request metadata.
Use case
For Agent Kit benchmarking, I may run the same task prompt against several profiles, branches, reruns, or criteria-focused variants. A tag such as purpose=test-vector-index-guidance or batch=agent-kit-split-vs-monolith-rerun would make it much easier to filter the Runs table and avoid accidentally mixing unrelated runs in aggregate results.
Acceptance criteria
- Request/run creation accepts optional key/value tags.
- The Runs table supports filtering by tag key and tag value.
- Request/run APIs return tags.
- Tags are preserved across retries/reruns or clearly copied to replacement requests.
- Tags are visible enough that users can tell "this run was to test X" without opening external notes.
- Dominant language
- TypeScript
- Stars
- 7
- Forks
- 13
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 26
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/scope
-
type: worker-update
Difficulty 1/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
-
type: worker-update
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
type: worker-update
Difficulty 1/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
author: JaGord documentation good first issue UI
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
author: cedricvidal bug portal
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
Similar issues
-
bug via-triage
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
pingdotgg/t3code#15221 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
521xueweihan/HelloGitHub#3847 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
MystenLabs/MemWal#1085 · 1 comment ·
Maintainers usually reply within 1 day
-
needs-triage
Difficulty 1/5 Under an hour Newbie friendliness 92/100
PhyberApex/kuroshiro#1187 ·
Maintainers usually reply within 1 day