Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Proposal: compare agent efficiency against verified task outcomes

オープン 初心者向け
#9 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
70/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
ai

調査の方向性

既存のAgentic AI Impact ExplorerとLLM推論の例を確認し、メンテナの優先する配置(拡張、追加、または単独実装)に合わせて調整します。指定された固定タスクセット、明示的な受け入れチェック、コスト追跡(成功例がない場合の成功あたり未定義コストを含む)、ラベル付き合成フィクスチャ、前提条件の注意事項を含む小さな再現可能な例を構築します。対象コンポーネントのテストを実行し、新しい例が正しく統合されることを確認します。

索引モデルが issue の本文から書いたものです。

説明

Hi,

I looked through the Agentic AI Impact Explorer and the LLM inference reference implementation. I would like to contribute a small example that makes task success explicit when comparing agent efficiency.

The explorer models resource use and retry overhead, while the inference example compares baseline and optimized prompts. A useful complement would show whether the lower-cost configuration still completes the same task to an agreed acceptance standard. Cost per attempt can improve even when cost per successful task gets worse.

I propose an offline, reproducible example with:

  • A fixed task set and explicit acceptance checks applied consistently to both configurations.
  • Observed task success, total cost across all attempts, and cost per successful task, with failed attempts included and retries counted once.
  • Clear separation between observed outcomes/cost data and any modeled energy or carbon values. No conversion from token count to electricity use without a stated estimation method.
  • A small synthetic fixture demonstrating the tradeoff, clearly labeled as illustrative rather than an empirical finding. A zero-success configuration would have undefined cost per success, not zero.
  • A short explanation of task boundaries, assumptions, uncertainty, and when the comparison is not meaningful.

I work on agent efficiency and evaluation and built TraceBurn, an open-source agent tracer and efficiency profiler: https://github.com/TommyTranX/traceburn.

Would you prefer this as an extension to the existing explorer, an addition to the inference example, or a separate community implementation? I can scope the contribution around the maintainers' preferred route before starting the implementation.

Hope it's useful.
Tommy

主要言語
Python
スター
9
フォーク
1
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Green-Software-Foundation/reference-implementations のほかの issue

Green-Software-Foundation/reference-implementations の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。