FEAT GUI: Expose built-in dataset provider metadata
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 45/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- ai, backend-api-design, data
Línea de trabajo
Start with SeedDatasetProvider, SeedDatasetMetadata, provider metadata, and the dataset backend service/models; review the identity/summary contract and catalog from items 1 and 3. Inspect existing name-list behavior and the dataset/Python/model/test instructions before tracing the explorer catalog. Done means read-only loaded, provider-only, and combined views preserve provenance, handle missing metadata safely, avoid remote loading, and include compatibility and metadata tests.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Is your feature request related to a problem? Please describe.
The explorer should help users discover built-in datasets that are not loaded yet. Provider names alone do not explain modalities, size, variants, sources, or whether credentials and gated access are required.
Work item 9 of 17 in #2744. Technical prerequisites: the dataset identity/summary contract and catalog from items 1 and 3. This starts the provider-loading milestone after item 8; keep not ready yet until promoted.
Describe the solution you'd like
Add metadata-only descriptors for registered built-in dataset providers and integrate them with the existing dataset catalog contract.
- Distinguish in-memory datasets from available providers and expose a loaded/available filter without confusing name collisions or partial/variant loads.
- Include available description, source URL, license/provenance, modality, harm categories, size information, and supported variants/options. Missing information remains unknown.
- Expose only documented, supported non-secret load options. Use registered provider identities; do not accept arbitrary Python paths, code, or unvalidated kwargs from clients.
- Describe credential and external gated-access prerequisites without returning credential values. A configured credential is not proof that its account has dataset access.
- Keep enumeration side-effect free: no dataset materialization, downloads, model calls, or credential-dependent fetches merely to display cards.
Acceptance criteria:
- Loaded-only, provider-only, and combined views distinguish provenance and availability correctly.
- Known provider metadata is displayed without confusing an estimate or coarse size category with an exact loaded count.
- Missing metadata or an unsupported load option does not break the whole catalog.
- Gated/public examples can be described without real credentials or remote fetches.
- Existing name-list clients remain compatible and metadata/read-only tests are included.
Describe alternatives you've considered, if relevant
Avoid a hardcoded frontend dataset list, a second provider registry, fetching every dataset for metadata, or a generic code-execution/loader endpoint. This issue does not load datasets; item 10 adds that operation.
Additional context
Start with SeedDatasetProvider, SeedDatasetMetadata, provider metadata, and the dataset backend service/models. Preserve the framework rule that downstream consumers read seeds from memory after loading. Coordinate count provenance with #2735 and follow dataset/Python/model/test instructions.
- Lenguaje dominante
- Python
- Estrellas
- 4.5k
- Forks
- 896
- Merge medio
- 3 d 1 h
- PR fusionados (30 d)
- 208
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/PyRIT
-
Multimodal writes fail entirely when embeddings are enabled (non-text piece crashes the add)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
Todos los issues de microsoft/PyRIT
Issues similares
-
good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
vllm-project/vllm-metal#822 ·
Los mantenedores suelen responder en 1 día
-
vector-store
Dificultad 1/5 1-3 horas Aptitud para principiantes 90/100
mem0ai/mem0#7461 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[Bug]: chunk_span_bounds and _validated_chunk_spans reject Pydantic models ChunkSpan and AudioFileAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
BasedHardware/omi#19047 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día