Proposal: compare agent efficiency against verified task outcomes
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 70/100
Research direction
Review the existing Agentic AI Impact Explorer and LLM inference example to align with the maintainers' preferred placement (extension, addition, or separate implementation). Build the small reproducible example with the specified fixed task set, explicit acceptance checks, cost tracking (including undefined cost per success for zero-success cases), labeled synthetic fixture, and assumptions note. Run tests for the target component to confirm the new example integrates correctly.
Written by the indexing model from the issue text.
Description
Hi,
I looked through the Agentic AI Impact Explorer and the LLM inference reference implementation. I would like to contribute a small example that makes task success explicit when comparing agent efficiency.
The explorer models resource use and retry overhead, while the inference example compares baseline and optimized prompts. A useful complement would show whether the lower-cost configuration still completes the same task to an agreed acceptance standard. Cost per attempt can improve even when cost per successful task gets worse.
I propose an offline, reproducible example with:
- A fixed task set and explicit acceptance checks applied consistently to both configurations.
- Observed task success, total cost across all attempts, and cost per successful task, with failed attempts included and retries counted once.
- Clear separation between observed outcomes/cost data and any modeled energy or carbon values. No conversion from token count to electricity use without a stated estimation method.
- A small synthetic fixture demonstrating the tradeoff, clearly labeled as illustrative rather than an empirical finding. A zero-success configuration would have undefined cost per success, not zero.
- A short explanation of task boundaries, assumptions, uncertainty, and when the comparison is not meaningful.
I work on agent efficiency and evaluation and built TraceBurn, an open-source agent tracer and efficiency profiler: https://github.com/TommyTranX/traceburn.
Would you prefer this as an extension to the existing explorer, an addition to the inference example, or a separate community implementation? I can scope the contribution around the maintainers' preferred route before starting the implementation.
Hope it's useful.
Tommy
- Dominant language
- Python
- Stars
- 9
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Green-Software-Foundation/reference-implementations
-
tools/community/sustainability-score, a repo/PR-level Sustainability Score reference implementationOpen
Difficulty 5/5 Over a week Newbie friendliness 25/100
All issues in Green-Software-Foundation/reference-implementations
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
pyjanitor-devs/pyjanitor#1758 ·
Maintainers usually reply within 1 day
-
bug ready for review
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
odysseus-dev/odysseus#6641 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
happypawspillaro/happypaws#78 ·
Maintainers usually reply within 4 days
-
pydanty:is-working
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
pydantic/pydantic-ai#10020 ·
Maintainers usually reply within 1 day
-
stdlib type-bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
python/cpython#159044 · 4 comments ·
Maintainers usually reply within 1 day