ObjectiveScorerEvaluator scores every conversation message as an assistant response
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 88/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
- 技術スタック
- python
- 領域
- ai, testing-qa
調査の方向性
pyrit/score/scorer_evaluation/scorer_evaluator.py の ObjectiveScorerEvaluator._validate_and_extract_data から始め、656-673 行目の harm ブランチと比較してください。既存のテスト test_validate_and_extract_harm_data_scores_only_assistant_message を読み、対応する objective のケースを追加してください。完了条件は、[user, assistant] エントリが正常に評価され、assistant メッセージがちょうど 1 つないエントリが名前によって拒否されることです。
索引モデルが issue の本文から書いたものです。
説明
Problem
HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.
HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).
ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.
Reproduction
Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):
evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0
The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.
Expected
The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.
Suggested fix
Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.
- 主要言語
- Python
- スター
- 4.5k
- フォーク
- 896
- 平均マージ
- 3日 2時間
- マージ済み PR(30日)
- 210
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
microsoft/PyRIT のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
microsoft/PyRIT#2888 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 91/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
Bug: triage GUI help wanted
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
microsoft/PyRIT#2868 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
microsoft/PyRIT の issue をすべて見る
似ている issue
-
docs pydanty:is-working
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
pydantic/pydantic-ai#8863 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
run-llama/llama_index#23278 ·
メンテナーはふだん 2 日以内に返信
-
documentation from-review-extraction github-actions priority: low severity:nit
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
LearningCircuit/local-deep-research#6946 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
oracle/langchain-oracle#323 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
tenstorrent/tt-metal#58057 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信