Labelled responses from remote datasets can't reach `HumanLabeledDataset`
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 48/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- huggingface, python
調査の方向性
datasets/seed_datasets/remote/aegis_ai_content_safety_dataset.py と HumanLabeledDataset のパスから始め、既存の HARM_CATEGORY_ALIAS_OVERRIDES マッピングとローダーの動作を datasets/scorer_evals/refusal.csv と比較します。Aegis 2.0 の限定的なケースに対する焦点を絞ったテストを追加し、assistant_response と選択した人手ラベルを保持した、使用可能な scorer_evals CSV を生成します。
索引モデルが issue の本文から書いたものです。
説明
There's no path from the remote dataset loaders to the labelled-data side of scoring.
What's there now. HumanLabeledDataset is only constructible via from_csv(), so scorer evaluation data has to be authored by hand. That's reflected in datasets/scorer_evals/ — refusal.csv is ~105 rows and several of the objective sets are single digits.
What's being dropped. datasets/seed_datasets/remote/ has 61 loaders, and several of the underlying HuggingFace datasets ship human labels alongside model responses. Those are discarded at load time because SeedDataset models prompts only. beaver_tails_dataset.py states it directly:
This loader extracts only the prompts (not the responses) and filters to unsafe entries by default.
aegis_ai_content_safety_dataset.py is the same shape. wildguardmix_dataset.py at least retains its classifier labels in metadata, but still produces a SeedDataset.
So labelled (response, score) pairs are already flowing through the loaders and being dropped, while the one class that needs them can only be fed by hand.
Proposal. A second output path off the existing remote loaders producing HumanLabeledDataset — reusing the fetch logic and the HARM_CATEGORY_ALIAS_OVERRIDES mapping that's already there, but retaining assistant_response and the human label.
Suggest starting narrow: Aegis 2.0 only (CC-BY-4.0, so no licence friction for a repo shipping under MIT) and one harm category that maps cleanly, with tests and a generated scorer_evals CSV so it's usable on merge.
Happy to build it if the shape is right — is HumanLabeledDataset the correct target here, or is there a reason the loaders are prompts-only that I'm missing?
- 主要言語
- Python
- スター
- 4.5k
- フォーク
- 896
- 平均マージ
- 2日 22時間
- マージ済み PR(30日)
- 214
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
microsoft/PyRIT のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 2 日以内に返信
-
Bug: triage GUI help wanted
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
microsoft/PyRIT#2868 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信
microsoft/PyRIT の issue をすべて見る
似ている issue
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedオープンworkflow
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
メンテナーはふだん 1 日以内に返信
-
metadata submission
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
canonical/content-cache-operator#163 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
[submission]オープンsubmission
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
leanprover/lean-eval-submissions#1852 ·
メンテナーはふだん 1 日以内に返信