A float-scale aggregate drops rationale-less constituents from its rationale
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 84/100
Research direction
Start in pyrit/score/float_scale/float_scale_score_aggregator.py, comparing the filtered rationale branch with _undetermined_result and true_false_score_aggregator._build_rationale. Check format_score_for_rationale and the FloatScaleThresholdScorer docstring, then add coverage for rationale-less constituents; done means every constituent is represented consistently in a multi-constituent float-scale rationale.
Written by the indexing model from the issue text.
Description
When a float-scale aggregate combines more than one constituent, the constituents without a rationale disappear from the rationale entirely.
In pyrit/score/float_scale/float_scale_score_aggregator.py:
else:
description = aggregate_description
# Only include scores with non-empty rationales
rationale_parts = [format_score_for_rationale(s) for s in scores if s.score_rationale]
rationale = "\n".join(rationale_parts) if rationale_parts else ""
Two things in the same package say the opposite. format_score_for_rationale — the formatter being called — is built to render a line for a rationale-less score (f" - {class_type} {value}: {score.score_rationale or ''}"), and its docstring describes the value and the rationale as what a line carries. The other two rationale builders do not filter: the undetermined branch of this same file (_undetermined_result) and true_false_score_aggregator._build_rationale both pass every constituent through.
The scorer that makes this visible is one this repository already documents: FloatScaleThresholdScorer's own docstring notes that AzureContentFilterScorer "routinely does not" supply a rationale. A multi-chunk or multi-category Azure filter aggregate therefore persists score_rationale == "" — the score is right, but nothing records what was aggregated or that there was more than one constituent, while the same run's true/false aggregates list theirs.
Proposal: drop the filter, so the float-scale rationale matches its two siblings. If the filter is deliberate, the alternative is to keep it and say so in the output (for example a trailing "N constituent(s) had no rationale"), so an empty rationale is distinguishable from a single-component one.
Happy to send the one-line change plus tests either way — I did not want to just delete a line that was written on purpose without asking.
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 896
- Avg merge
- 2d 16h
- Merged PRs (30d)
- 230
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/PyRIT
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
microsoft/PyRIT#2948 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 2 days
-
Bug: triage GUI help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
microsoft/PyRIT#2868 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
-
Bug: triage
Difficulty 3/5 1-2 days Newbie friendliness 57/100
Maintainers usually reply within 2 days
Similar issues
-
upstream update
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
conan-io/conan-center-index#31098 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
john-kurkowski/tldextract#382 ·
-
comp/tools duplicate P2 sweeper:risk-compatibility tool/mcp type/bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
NousResearch/hermes-agent#132042 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#13092 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 85/100
feder-cr/invisible_playwright_mcp#1408 ·
Maintainers usually reply within 1 day