Proposal: compare agent efficiency against verified task outcomes
まだ誰も着手していません。
評価
調査の方向性
既存のAgentic AI Impact ExplorerとLLM推論の例を確認し、メンテナの優先する配置(拡張、追加、または単独実装)に合わせて調整します。指定された固定タスクセット、明示的な受け入れチェック、コスト追跡(成功例がない場合の成功あたり未定義コストを含む)、ラベル付き合成フィクスチャ、前提条件の注意事項を含む小さな再現可能な例を構築します。対象コンポーネントのテストを実行し、新しい例が正しく統合されることを確認します。
索引モデルが issue の本文から書いたものです。
説明
Hi,
I looked through the Agentic AI Impact Explorer and the LLM inference reference implementation. I would like to contribute a small example that makes task success explicit when comparing agent efficiency.
The explorer models resource use and retry overhead, while the inference example compares baseline and optimized prompts. A useful complement would show whether the lower-cost configuration still completes the same task to an agreed acceptance standard. Cost per attempt can improve even when cost per successful task gets worse.
I propose an offline, reproducible example with:
- A fixed task set and explicit acceptance checks applied consistently to both configurations.
- Observed task success, total cost across all attempts, and cost per successful task, with failed attempts included and retries counted once.
- Clear separation between observed outcomes/cost data and any modeled energy or carbon values. No conversion from token count to electricity use without a stated estimation method.
- A small synthetic fixture demonstrating the tradeoff, clearly labeled as illustrative rather than an empirical finding. A zero-success configuration would have undefined cost per success, not zero.
- A short explanation of task boundaries, assumptions, uncertainty, and when the comparison is not meaningful.
I work on agent efficiency and evaluation and built TraceBurn, an open-source agent tracer and efficiency profiler: https://github.com/TommyTranX/traceburn.
Would you prefer this as an extension to the existing explorer, an addition to the inference example, or a separate community implementation? I can scope the contribution around the maintainers' preferred route before starting the implementation.
Hope it's useful.
Tommy
- 主要言語
- Python
- スター
- 9
- フォーク
- 1
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Green-Software-Foundation/reference-implementations のほかの issue
-
tools/community/sustainability-score, a repo/PR-level Sustainability Score reference implementationオープン
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
Green-Software-Foundation/reference-implementations の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
SR_SECURITY_DESCRIPTOR.fromString drops the SACL when no DACL is present対応中かも @paul7436 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
equinor/fmu-sumo-uploader#302 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
modelscope/evalscope#1821 ·
メンテナーはふだん 1 日以内に返信
-
Sanity on ansible-core devel fails: ignore-2.23.txt references the removed import-3.9 test対応中かも @yurnov が今日担当しました。 オープンneeds_triage
難易度 1/5 1時間未満 初心者へのやさしさ 91/100
ansible-collections/kubernetes.core#1275 ·
メンテナーはふだん 1 日以内に返信