Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

ObjectiveScorerEvaluator scores every conversation message as an assistant response

オープン 初心者向け
#2,835 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
88/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
ai, testing-qa

調査の方向性

pyrit/score/scorer_evaluation/scorer_evaluator.py の ObjectiveScorerEvaluator._validate_and_extract_data から始め、656-673 行目の harm ブランチと比較してください。既存のテスト test_validate_and_extract_harm_data_scores_only_assistant_message を読み、対応する objective のケースを追加してください。完了条件は、[user, assistant] エントリが正常に評価され、assistant メッセージがちょうど 1 つないエントリが名前によって拒否されることです。

索引モデルが issue の本文から書いたものです。

説明

Problem

HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.

HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).

ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.

Reproduction

Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):

evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0

The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.

Expected

The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.

Suggested fix

Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.

主要言語
Python
スター
4.5k
フォーク
896
平均マージ
3日 2時間
マージ済み PR(30日)
210

環境構築

このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

microsoft/PyRIT のほかの issue

microsoft/PyRIT の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。