Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Community package: azure-ai-evaluation-openeval-adapter (evaluate() data/results <-> EvalPort interchange)

未关闭
#48,971 0 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
1/5
预计耗时
1 小时以内
新手友好度
35/100
Issue 类型
文档
描述清晰度
需要澄清
活跃度
活跃
技术栈
azure, python
领域
documentation

调研方向

从链接的 azure-ai-evaluation-openeval-adapter README 开始,检查此 repository 是否有 community-packages 列表或其他文档入口。不要求修改 repository;如果存在合适的列表,done 将是一个简洁的发现性链接,否则需要澄清 issue。

由索引模型根据 Issue 内容生成。

描述

Evaluation feature-request

Hi Azure AI Evaluation team — I maintain EvalPort (Apache 2.0), an open, schema-validated JSON interchange format for portable LLM evaluation test cases, graders, suites, and results. It already has independently-tested adapter packages for ~20 other eval/observability frameworks (MLflow, LangSmith, Ragas, Vertex AI Gen AI Evaluation, Hugging Face evaluate, and others), and I built one for azure-ai-evaluation the same way.

This isn't a request for a change in this repo — I'm not proposing new API surface or asking for a design review, just flagging a working, tested community package in case it's useful to know about or link from docs.

azure-ai-evaluation-openeval-adapter

from azure.ai.evaluation import F1ScoreEvaluator, evaluate
from azure_ai_evaluation_openeval_adapter import to_openeval, evaluation_result_to_openeval
from openeval.validate import validate_suite, validate_result_set

suite = to_openeval(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()}, suite_id="my_eval_suite")
assert validate_suite(suite).valid

result = evaluate(data="my_eval_data.jsonl", evaluators={"f1": F1ScoreEvaluator()})
result_set = evaluation_result_to_openeval(result, suite_id="my_eval_suite")
assert validate_result_set(result_set).valid

to_openeval() accepts exactly what evaluate() itself accepts for data/evaluators, so it's a pure format bridge rather than new infrastructure. The one design choice worth flagging: every evaluator (local NLP metrics like F1/BLEU/ROUGE, AI-assisted evaluators needing a live model_config, and the content-safety evaluators needing a live Foundry project) maps to EvalPort's custom grader type rather than being force-fit into semantic_similarity or llm_judge — those types require params (threshold, prompt) this adapter can't honestly fabricate from the outside. Full mapping table and the flat-row parsing logic (recovering per-metric score/passed/reason from evaluate()'s real outputs.<evaluator>.* column convention) are in the README.

21 tests, all passing locally against the real installed azure-ai-evaluation package and EvalPort's real validate_suite()/validate_result_set() — not mocked.

No action needed — this lives entirely outside azure-sdk-for-python as an independent package (pip install via git+, not yet on PyPI). Flagging mainly for discoverability; happy to adjust the mapping if the evaluation module's public API shifts, or to send a one-line docs PR if there's a community-packages list this belongs on.

Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md

主要语言
Python
星标
5.6k
派生
3.4k
平均合并
2 天 1 分钟
30 天内合并 PR
218

环境准备

在 Codespaces 中打开

在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

Azure/azure-sdk-for-python 的其他 Issue

查看 Azure/azure-sdk-for-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。