Community package: azure-ai-evaluation-openeval-adapter (evaluate() data/results <-> EvalPort interchange)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 1/5
- Tiempo estimado
- Menos de una hora
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Documentación
- Claridad
- Necesita aclaración
- Estado de actividad
- Activo
- Stack tecnológico
- azure, python
- Área
- documentation
Línea de trabajo
Empieza por el README enlazado de azure-ai-evaluation-openeval-adapter y comprueba si este repositorio tiene una lista de community-packages u otro punto de entrada a la documentación. No se solicita ningún cambio en el repositorio; si existe una lista adecuada, done sería un enlace conciso para mejorar su descubribilidad; de lo contrario, es necesario aclarar el issue.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Hi Azure AI Evaluation team — I maintain EvalPort (Apache 2.0), an open, schema-validated JSON interchange format for portable LLM evaluation test cases, graders, suites, and results. It already has independently-tested adapter packages for ~20 other eval/observability frameworks (MLflow, LangSmith, Ragas, Vertex AI Gen AI Evaluation, Hugging Face evaluate, and others), and I built one for azure-ai-evaluation the same way.
This isn't a request for a change in this repo — I'm not proposing new API surface or asking for a design review, just flagging a working, tested community package in case it's useful to know about or link from docs.
azure-ai-evaluation-openeval-adapter
from azure.ai.evaluation import F1ScoreEvaluator, evaluate
from azure_ai_evaluation_openeval_adapter import to_openeval, evaluation_result_to_openeval
from openeval.validate import validate_suite, validate_result_set
suite = to_openeval(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()}, suite_id="my_eval_suite")
assert validate_suite(suite).valid
result = evaluate(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()})
result_set = evaluation_result_to_openeval(result, suite_id="my_eval_suite")
assert validate_result_set(result_set).valid
to_openeval() accepts exactly what evaluate() itself accepts for data/evaluators, so it's a pure format bridge rather than new infrastructure. The one design choice worth flagging: every evaluator (local NLP metrics like F1/BLEU/ROUGE, AI-assisted evaluators needing a live model_config, and the content-safety evaluators needing a live Foundry project) maps to EvalPort's custom grader type rather than being force-fit into semantic_similarity or llm_judge — those types require params (threshold, prompt) this adapter can't honestly fabricate from the outside. Full mapping table and the flat-row parsing logic (recovering per-metric score/passed/reason from evaluate()'s real outputs.<evaluator>.* column convention) are in the README.
21 tests, all passing locally against the real installed azure-ai-evaluation package and EvalPort's real validate_suite()/validate_result_set() — not mocked.
No action needed — this lives entirely outside azure-sdk-for-python as an independent package (pip install via git+, not yet on PyPI). Flagging mainly for discoverability; happy to adjust the mapping if the evaluation module's public API shifts, or to send a one-line docs PR if there's a community-packages list this belongs on.
Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md
- Lenguaje dominante
- Python
- Estrellas
- 5.6k
- Forks
- 3.4k
- Merge medio
- 1 d 18 h
- PR fusionados (30 d)
- 202
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Azure/azure-sdk-for-python
-
Evaluation Service Attention
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Azure/azure-sdk-for-python#49190 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Update CODEOWNERSAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
Azure/azure-sdk-for-python#49183 · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Evaluation Service Attention
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Azure/azure-sdk-for-python#49153 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Search Service Attention
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
Azure/azure-sdk-for-python#48555 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Azure.Core customer-reported feature-request needs-team-attention
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Azure/azure-sdk-for-python#47186 ·
Los mantenedores suelen responder en 1 día
Todos los issues de Azure/azure-sdk-for-python
Issues similares
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedAbiertoworkflow
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
Los mantenedores suelen responder en 1 día
-
New Submission: TropWATERAbiertometadata submission
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
-
Wrongly named dashboard variableAbiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
canonical/content-cache-operator#163 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[submission]Abiertosubmission
Dificultad 1/5 Menos de una hora Aptitud para principiantes 65/100
leanprover/lean-eval-submissions#1852 ·
Los mantenedores suelen responder en 1 día