Labelled responses from remote datasets can't reach `HumanLabeledDataset`
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- huggingface, python
- Área
- data, machine-learning, testing-qa
Línea de trabajo
Comienza con datasets/seed_datasets/remote/aegis_ai_content_safety_dataset.py y la ruta de HumanLabeledDataset; después, compara el mapeo HARM_CATEGORY_ALIAS_OVERRIDES existente y el comportamiento del cargador con datasets/scorer_evals/refusal.csv. Añade pruebas específicas para el caso concreto de Aegis 2.0 y genera un CSV utilizable de scorer_evals que conserve assistant_response y la etiqueta humana seleccionada.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
There's no path from the remote dataset loaders to the labelled-data side of scoring.
What's there now. HumanLabeledDataset is only constructible via from_csv(), so scorer evaluation data has to be authored by hand. That's reflected in datasets/scorer_evals/ — refusal.csv is ~105 rows and several of the objective sets are single digits.
What's being dropped. datasets/seed_datasets/remote/ has 61 loaders, and several of the underlying HuggingFace datasets ship human labels alongside model responses. Those are discarded at load time because SeedDataset models prompts only. beaver_tails_dataset.py states it directly:
This loader extracts only the prompts (not the responses) and filters to unsafe entries by default.
aegis_ai_content_safety_dataset.py is the same shape. wildguardmix_dataset.py at least retains its classifier labels in metadata, but still produces a SeedDataset.
So labelled (response, score) pairs are already flowing through the loaders and being dropped, while the one class that needs them can only be fed by hand.
Proposal. A second output path off the existing remote loaders producing HumanLabeledDataset — reusing the fetch logic and the HARM_CATEGORY_ALIAS_OVERRIDES mapping that's already there, but retaining assistant_response and the human label.
Suggest starting narrow: Aegis 2.0 only (CC-BY-4.0, so no licence friction for a repo shipping under MIT) and one harm category that maps cleanly, with tests and a generated scorer_evals CSV so it's usable on merge.
Happy to build it if the shape is right — is HumanLabeledDataset the correct target here, or is there a reason the loaders are prompts-only that I'm missing?
- Lenguaje dominante
- Python
- Estrellas
- 4.5k
- Forks
- 896
- Merge medio
- 2 d 22 h
- PR fusionados (30 d)
- 214
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/PyRIT
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
microsoft/PyRIT#2905 · 3 comentarios ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 2 días
-
Bug: triage GUI help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
microsoft/PyRIT#2868 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 2 días
-
feature-request
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
Los mantenedores suelen responder en 2 días
Todos los issues de microsoft/PyRIT
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
Los mantenedores suelen responder en 1 día
-
https://search.utilibre.orgAbiertoinstance instance add
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
searxng/searx-instances#941 · 1 comentario ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día