Community package: azure-ai-evaluation-openeval-adapter (evaluate() data/results <-> EvalPort interchange)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 1/5
- Tempo stimato
- Meno di un'ora
- Idoneità per principianti
- 35/100
- Tipo di issue
- Documentazione
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- azure, python
- Ambito
- documentation
Direzione di ricerca
Inizia dal README collegato di azure-ai-evaluation-openeval-adapter e verifica se questo repository dispone di un elenco di community-packages o di un altro punto di accesso alla documentazione. Non è richiesta alcuna modifica al repository; se esiste un elenco appropriato, done sarebbe un link conciso per facilitarne l’individuazione, altrimenti l’issue deve essere chiarita.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi Azure AI Evaluation team — I maintain EvalPort (Apache 2.0), an open, schema-validated JSON interchange format for portable LLM evaluation test cases, graders, suites, and results. It already has independently-tested adapter packages for ~20 other eval/observability frameworks (MLflow, LangSmith, Ragas, Vertex AI Gen AI Evaluation, Hugging Face evaluate, and others), and I built one for azure-ai-evaluation the same way.
This isn't a request for a change in this repo — I'm not proposing new API surface or asking for a design review, just flagging a working, tested community package in case it's useful to know about or link from docs.
azure-ai-evaluation-openeval-adapter
from azure.ai.evaluation import F1ScoreEvaluator, evaluate
from azure_ai_evaluation_openeval_adapter import to_openeval, evaluation_result_to_openeval
from openeval.validate import validate_suite, validate_result_set
suite = to_openeval(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()}, suite_id="my_eval_suite")
assert validate_suite(suite).valid
result = evaluate(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()})
result_set = evaluation_result_to_openeval(result, suite_id="my_eval_suite")
assert validate_result_set(result_set).valid
to_openeval() accepts exactly what evaluate() itself accepts for data/evaluators, so it's a pure format bridge rather than new infrastructure. The one design choice worth flagging: every evaluator (local NLP metrics like F1/BLEU/ROUGE, AI-assisted evaluators needing a live model_config, and the content-safety evaluators needing a live Foundry project) maps to EvalPort's custom grader type rather than being force-fit into semantic_similarity or llm_judge — those types require params (threshold, prompt) this adapter can't honestly fabricate from the outside. Full mapping table and the flat-row parsing logic (recovering per-metric score/passed/reason from evaluate()'s real outputs.<evaluator>.* column convention) are in the README.
21 tests, all passing locally against the real installed azure-ai-evaluation package and EvalPort's real validate_suite()/validate_result_set() — not mocked.
No action needed — this lives entirely outside azure-sdk-for-python as an independent package (pip install via git+, not yet on PyPI). Flagging mainly for discoverability; happy to adjust the mapping if the evaluation module's public API shifts, or to send a one-line docs PR if there's a community-packages list this belongs on.
Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md
- Lingua principale
- Python
- Stelle
- 5.6k
- Fork
- 3.4k
- Merge medio
- 1g 18h
- PR unite (30g)
- 202
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Azure/azure-sdk-for-python
-
Evaluation Service Attention
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
Azure/azure-sdk-for-python#49190 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Update CODEOWNERSAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
Azure/azure-sdk-for-python#49183 · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Evaluation Service Attention
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Azure/azure-sdk-for-python#49153 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Search Service Attention
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
Azure/azure-sdk-for-python#48555 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Azure.Core customer-reported feature-request needs-team-attention
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Azure/azure-sdk-for-python#47186 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di Azure/azure-sdk-for-python
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
PedestrianDynamics/pyFDS-Evac#343 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
theskumar/python-dotenv#708 ·
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 2 giorni
-
Docs Timedelta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
pandas-dev/pandas#69919 ·
I maintainer di solito rispondono entro 1 giorno
-
API documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
zephyrproject-rtos/west#1009 · 2 commenti ·
I maintainer di solito rispondono entro 3 giorni