ObjectiveScorerEvaluator scores every conversation message as an assistant response
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 88/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- ai, testing-qa
Research direction
Start in pyrit/score/scorer_evaluation/scorer_evaluator.py at ObjectiveScorerEvaluator._validate_and_extract_data, then compare it with the harm branch at lines 656-673. Read the existing test_validate_and_extract_harm_data_scores_only_assistant_message test and add the corresponding objective cases; done means [user, assistant] entries evaluate successfully while entries without exactly one assistant message are rejected by name.
Written by the indexing model from the issue text.
Description
Problem
HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.
HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).
ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.
Reproduction
Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):
evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0
The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.
Expected
The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.
Suggested fix
Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 896
- Avg merge
- 3d 2h
- Merged PRs (30d)
- 210
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/PyRIT
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
microsoft/PyRIT#2888 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 91/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
Bug: triage GUI help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
microsoft/PyRIT#2868 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
docs pydanty:is-working
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
pydantic/pydantic-ai#8863 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
run-llama/llama_index#23278 ·
Maintainers usually reply within 2 days
-
documentation from-review-extraction github-actions priority: low severity:nit
Difficulty 1/5 Under an hour Newbie friendliness 92/100
LearningCircuit/local-deep-research#6946 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
oracle/langchain-oracle#323 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
tenstorrent/tt-metal#58057 · 1 comment ·
Maintainers usually reply within 1 day