ObjectiveScorerEvaluator scores every conversation message as an assistant response
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 88/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- ai, testing-qa
Hướng nghiên cứu
Bắt đầu trong pyrit/score/scorer_evaluation/scorer_evaluator.py tại ObjectiveScorerEvaluator._validate_and_extract_data, sau đó so sánh với nhánh harm ở các dòng 656-673. Đọc test hiện có test_validate_and_extract_harm_data_scores_only_assistant_message và thêm các trường hợp objective tương ứng; hoàn thành khi các mục [user, assistant] được đánh giá thành công, còn các mục không có chính xác một thông báo assistant bị từ chối theo tên.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Problem
HumanLabeledEntry.conversation is documented as "a list of Message objects representing the conversation to be scored. This can contain one Message object if you are just scoring individual assistant responses" (pyrit/score/scorer_evaluation/human_labeled_dataset.py:30-36), so a [user, assistant] conversation is a supported shape.
HarmScorerEvaluator handles that shape: every message is seeded into memory, only the assistant message is handed to the scorer, and an entry without exactly one assistant message is rejected by name (pyrit/score/scorer_evaluation/scorer_evaluator.py:656-673, pinned by test_validate_and_extract_harm_data_scores_only_assistant_message).
ObjectiveScorerEvaluator._validate_and_extract_data does not. pyrit/score/scorer_evaluation/scorer_evaluator.py:783-789 appends every message of the conversation to assistant_responses, while human_scores_list and objectives each get one row per entry.
Reproduction
Two ObjectiveHumanLabeledEntry objects whose conversations are [user, assistant], a mocked TrueFalseScorer, then evaluate_dataset_async(...):
evaluate_dataset_async raised: ValueError: All arguments must have the same length.
score_async calls: 0
The run cannot start, and the message points nowhere useful: the mismatch is between the responses and the score rows built by the evaluator, not in the scorer. Single-message entries (what from_csv produces) are unaffected, so this only shows up once a dataset is built in code with the surrounding turn included — which is exactly what the entry docstring describes.
Expected
The same contract as the harm path: seed all messages into memory, score the assistant message, and name the entry when it does not hold exactly one assistant message.
Suggested fix
Mirror the harm branch in ObjectiveScorerEvaluator._validate_and_extract_data (about 8 lines) and add tests modelled on the existing harm one.
- Ngôn ngữ chính
- Python
- Star
- 4.5k
- Fork
- 896
- Merge trung bình
- 2 ngày 22 giờ
- Pull request đã merge (30 ngày)
- 220
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/PyRIT
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
microsoft/PyRIT#2905 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Bug: triage GUI help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
microsoft/PyRIT#2868 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
feature-request
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 30/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của microsoft/PyRIT
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
DOI-USGS/pywatershed#421 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
python-pillow/Pillow#10087 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày