Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

ObjectiveScorerEvaluator scores every conversation message as an assistant response

Đang mở Phù hợp với người mới
#2,835 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
88/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
ai, testing-qa

Hướng nghiên cứu

Bắt đầu trong pyrit/score/scorer_evaluation/scorer_evaluator.py tại ObjectiveScorerEvaluator._validate_and_extract_data, sau đó so sánh với nhánh harm ở các dòng 656-673. Đọc test hiện có test_validate_and_extract_harm_data_scores_only_assistant_message và thêm các trường hợp objective tương ứng; hoàn thành khi các mục [user, assistant] được đánh giá thành công, còn các mục không có chính xác một thông báo assistant bị từ chối theo tên.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Problem

HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.

HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).

ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.

Reproduction

Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):

evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0

The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.

Expected

The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.

Suggested fix

Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.

Ngôn ngữ chính
Python
Star
4.5k
Fork
896
Merge trung bình
2 ngày 22 giờ
Pull request đã merge (30 ngày)
220

Chuẩn bị môi trường

Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/PyRIT

Tất cả issue của microsoft/PyRIT

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.