Labelled responses from remote datasets can't reach `HumanLabeledDataset`
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- huggingface, python
- Domain
- data, machine-learning, testing-qa
Research direction
Start with datasets/seed_datasets/remote/aegis_ai_content_safety_dataset.py and the HumanLabeledDataset path, then compare the existing HARM_CATEGORY_ALIAS_OVERRIDES mapping and loader behavior with datasets/scorer_evals/refusal.csv. Add focused tests for the narrow Aegis 2.0 case and generate a usable scorer_evals CSV retaining assistant_response and the selected human label.
Written by the indexing model from the issue text.
Description
There's no path from the remote dataset loaders to the labelled-data side of scoring.
What's there now. HumanLabeledDataset is only constructible via from_csv(), so scorer evaluation data has to be authored by hand. That's reflected in datasets/scorer_evals/ — refusal.csv is ~105 rows and several of the objective sets are single digits.
What's being dropped. datasets/seed_datasets/remote/ has 61 loaders, and several of the underlying HuggingFace datasets ship human labels alongside model responses. Those are discarded at load time because SeedDataset models prompts only. beaver_tails_dataset.py states it directly:
This loader extracts only the prompts (not the responses) and filters to unsafe entries by default.
aegis_ai_content_safety_dataset.py is the same shape. wildguardmix_dataset.py at least retains its classifier labels in metadata, but still produces a SeedDataset.
So labelled (response, score) pairs are already flowing through the loaders and being dropped, while the one class that needs them can only be fed by hand.
Proposal. A second output path off the existing remote loaders producing HumanLabeledDataset — reusing the fetch logic and the HARM_CATEGORY_ALIAS_OVERRIDES mapping that's already there, but retaining assistant_response and the human label.
Suggest starting narrow: Aegis 2.0 only (CC-BY-4.0, so no licence friction for a repo shipping under MIT) and one harm category that maps cleanly, with tests and a generated scorer_evals CSV so it's usable on merge.
Happy to build it if the shape is right — is HumanLabeledDataset the correct target here, or is there a reason the loaders are prompts-only that I'm missing?
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 896
- Avg merge
- 2d 19h
- Merged PRs (30d)
- 234
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/PyRIT
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
microsoft/PyRIT#2948 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 2 days
-
Bug: triage GUI help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
microsoft/PyRIT#2868 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
-
Bug: triage
Difficulty 3/5 1-2 days Newbie friendliness 57/100
Maintainers usually reply within 2 days
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
raullenchai/Rapid-MLX#4042 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
LearningCircuit/local-deep-research#7067 ·
Maintainers usually reply within 1 day
-
#bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
apache/superset#44923 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
lawndoc/stack-back#123 ·