Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

DOC ConversationScorer comments say tool output is excluded from the scored text; the code includes it

已关闭
#2,706 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
50/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
python
领域
ai, security

调研方向

阅读 pyrit/score/conversation_scorer.py 和 pyrit/score/scorer_prompt_validator.py,然后运行 tests/unit/score/test_conversation_history_scorer.py 中相邻的测试,并添加一个 tool-role 片段。与 maintainers 确认是否应对 tool 输出进行评分,并根据该决定调整实现、注释和回归覆盖范围。

由索引模型根据 Issue 内容生成。

描述

Describe the bug

ConversationScorer builds the text that a harm/objective scorer judges, and three consecutive
comments in that function say raw tool output is deliberately left out of it — but the code includes
it, so a role="tool" piece is folded into the scored conversation and rendered as "Tool: …".

pyrit/score/conversation_scorer.py on main:

# Goes through each message in the conversation and appends user/assistant messages only
# Explicitly excludes system, tool, developer messages from being scored/included in conversation history
# they are allowed in validation but not included in the scored conversation text
for conv_message in conversation:
    for piece in conv_message.message_pieces:
        # Only include user and assistant messages in the conversation text
        if piece.api_role in ["user", "assistant", "tool"] and self._validator.is_role_supported(piece):

The validator does not close the gap: ScorerPromptValidator's supported_roles "Defaults to every
role except simulated_assistant" (pyrit/score/scorer_prompt_validator.py:45-47), so under the
default validator a tool piece passes both conditions.

Steps/Code to Reproduce

Offline unit-style reproduction (no keys, no network), following
tests/unit/score/test_conversation_history_scorer.py::test_conversation_history_scorer_filters_roles_correctly
and printing what the wrapped scorer actually receives. Pieces stored for one conversation id:
user, tool, developer, system, assistant, each with a distinct sentinel string; then
capture mock_scorer._score_nested_async.call_args.kwargs["scorable"].value.

Measured on microsoft/PyRIT main @ 2215c5b (Python 3.14.5, uv sync default groups):

SCORED TEXT repr:
'User: USER_SENTINEL\nTool: TOOL_SENTINEL\nAssistant: ASSISTANT_SENTINEL\n'
USER_SENTINEL      in scored text -> True
TOOL_SENTINEL      in scored text -> True
DEV_SENTINEL       in scored text -> False
SYS_SENTINEL       in scored text -> False
ASSISTANT_SENTINEL in scored text -> True

The neighbouring test already pins the intended shape for two of the roles — it asserts
expected_conversation = "User: User message\nAssistant: Assistant message\n" and
assert "System message" not in called_scorable.value — and it does not cover a tool piece at all,
which is why the divergence is invisible to CI.

Expected Results

One of the two, and this is the question I cannot answer from the code alone:

  • either tool output is excluded, matching the three comments (drop "tool" from the role list), or
  • tool output is included on purpose, in which case the comments are wrong and a test should pin the
    behaviour so the two cannot drift again.
Actual Results

Tool text is included in the conversation the scorer judges, while developer and system are not.

Why I think it matters for scoring correctness

A scorer is being asked to judge what the target produced. Third-party tool results are neither a
user turn nor an assistant turn, so a long or hostile tool payload sitting in the judged text can move
the verdict in either direction without reflecting the target's own behaviour. scorer_prompt_validator.py:9-14
gives the analogous reason for dropping simulated history — a scorer that judges the target must not
mistake it for what the target said.

The history explains the split rather than settling it: the comments came in with the original
ConversationScorer (#1138, 2337fce7, 2025-12-12), while "tool" was added to the whitelist by
#2518 "Policy Scorer Compatibility (phase 2.5)" (8b3826be, 2026-09-02), which left the adjacent
comments untouched. So I cannot tell whether #2518 intended policy scorers to see tool output and the
comments simply went stale, or whether the widening was incidental.

If tool visibility is required only for the policy-scorer path, a validator-driven option is to keep
the exclusion in ConversationScorer and let a scorer that wants tool output declare it through
supported_roles, so the include/exclude decision lives in one place instead of two.

Not verified, and I would not want a fix to assume it: I have not checked whether any shipped scenario
actually scores a conversation that contains tool pieces, so I cannot say how often this changes a real
verdict in practice — only that the default path includes them.

Versions
  • OS: macOS 27.0 (arm64)
  • Python: 3.14.5 (also reproduced logic review on 3.11.15)
  • PyRIT: 1.2.0.dev0 from main @ 2215c5b
主要语言
Python
星标
4.5k
派生
896
平均合并
2 天 19 小时
30 天内合并 PR
206

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/PyRIT 的其他 Issue

查看 microsoft/PyRIT 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。