ObjectiveScorerEvaluator scores every conversation message as an assistant response
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 88/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
- 技术栈
- python
- 领域
- ai, testing-qa
调研方向
从 pyrit/score/scorer_evaluation/scorer_evaluator.py 中的 ObjectiveScorerEvaluator._validate_and_extract_data 开始,然后将其与第 656-673 行的 harm 分支进行比较。阅读现有的 test_validate_and_extract_harm_data_scores_only_assistant_message 测试,并添加相应的 objective 用例;完成标准是 [user, assistant] 条目能够成功评估,而不恰好包含一条 assistant 消息的条目会按名称被拒绝。
由索引模型根据 Issue 内容生成。
描述
Problem
HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.
HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).
ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.
Reproduction
Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):
evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0
The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.
Expected
The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.
Suggested fix
Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.
- 主要语言
- Python
- 星标
- 4.5k
- 派生
- 896
- 平均合并
- 2 天 22 小时
- 30 天内合并 PR
- 214
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/PyRIT 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
microsoft/PyRIT#2905 · 3 条评论 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 2 天内回复
-
Bug: triage GUI help wanted
难度 2/5 1-3 小时 新手友好度 86/100
microsoft/PyRIT#2868 · 1 条评论 ·
维护者通常 2 天内回复
-
feature-request
难度 5/5 一周以上 新手友好度 30/100
维护者通常 2 天内回复
-
难度 3/5 1-2 天 新手友好度 72/100
维护者通常 2 天内回复
相似的 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 90/100
维护者通常 1 天内回复
-
instance instance add
难度 2/5 1-3 小时 新手友好度 68/100
searxng/searx-instances#941 · 1 条评论 ·
-
难度 1/5 1 小时以内 新手友好度 92/100
FluidNumerics/fluid-walk-blocker#89 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复