Proposal: compare agent efficiency against verified task outcomes
还没有人认领这个 Issue。
评估
调研方向
审查现有的Agentic AI Impact Explorer和LLM推理示例,以与维护者偏好的部署位置(扩展、附加或独立实现)保持一致。构建包含指定固定任务集、明确验收检查、成本跟踪(包括零成功案例的每成功次未定义成本)、带标签的合成测试夹具以及假设说明的小型可复现示例。运行目标组件的测试,确认新示例已正确集成。
由索引模型根据 Issue 内容生成。
描述
Hi,
I looked through the Agentic AI Impact Explorer and the LLM inference reference implementation. I would like to contribute a small example that makes task success explicit when comparing agent efficiency.
The explorer models resource use and retry overhead, while the inference example compares baseline and optimized prompts. A useful complement would show whether the lower-cost configuration still completes the same task to an agreed acceptance standard. Cost per attempt can improve even when cost per successful task gets worse.
I propose an offline, reproducible example with:
- A fixed task set and explicit acceptance checks applied consistently to both configurations.
- Observed task success, total cost across all attempts, and cost per successful task, with failed attempts included and retries counted once.
- Clear separation between observed outcomes/cost data and any modeled energy or carbon values. No conversion from token count to electricity use without a stated estimation method.
- A small synthetic fixture demonstrating the tradeoff, clearly labeled as illustrative rather than an empirical finding. A zero-success configuration would have undefined cost per success, not zero.
- A short explanation of task boundaries, assumptions, uncertainty, and when the comparison is not meaningful.
I work on agent efficiency and evaluation and built TraceBurn, an open-source agent tracer and efficiency profiler: https://github.com/TommyTranX/traceburn.
Would you prefer this as an extension to the existing explorer, an addition to the inference example, or a separate community implementation? I can scope the contribution around the maintainers' preferred route before starting the implementation.
Hope it's useful.
Tommy
- 主要语言
- Python
- 星标
- 9
- 派生
- 1
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
Green-Software-Foundation/reference-implementations 的其他 Issue
-
tools/community/sustainability-score, a repo/PR-level Sustainability Score reference implementation未关闭
难度 5/5 一周以上 新手友好度 25/100
查看 Green-Software-Foundation/reference-implementations 的全部 Issue
相似的 Issue
-
HTML: <template> content is extracted as document text可能已有人在做 @ryanmeowy 今天认领。 未关闭bug html
难度 1/5 1 小时以内 新手友好度 82/100
docling-project/docling#4714 · 2 条评论 ·
维护者通常 1 天内回复
-
[BUG] Qdrant RAG client applies score_threshold to raw cosine similarity, not the 0-1 score it returns可能已有人在做 @roydonsequeira 今天认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
ashhart/TensorFold#536 ·
维护者通常 1 天内回复
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
难度 1/5 1 小时以内 新手友好度 62/100
NousResearch/hermes-agent#135440 · 1 条评论 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 85/100
Deepak3699/Ai_Mentor#244 ·
维护者通常 1 天内回复