Backtest/eval harness for the LLM dispute bot (llm-bot)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 48/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- javascript
- 领域
- blockchain, testing-qa
调研方向
First confirm with a maintainer whether packages/llm-bot is still maintained. Then inspect Backtest.test, DisputerStrategy.process, and src/core/OptimisticOracleV2.ts under the existing hardhat test setup; use the resolved-request fixture and strategy interface described in the issue as the scope. Done means a tested batch utility reports accuracy, recall, and precision and CI detects regressions without live network or model calls.
由索引模型根据 Issue 内容生成。
描述
Why
The packages/llm-bot package is meant to automate a watchdog role on UMA's Optimistic Oracle: watch for proposed answers and use an LLM to decide whether a proposal looks wrong and should be disputed. Today that decision logic in the public repo is a placeholder rather than a working evaluator:
DisputerStrategy.process(src/core/OptimisticOracleV2.ts) doesn't call any model. It returns a fixed, hardcoded verdict every time (always "the correct answer is 1, and yes, dispute it").Backtest.testcan only check one request at a time against its final resolved value.OptimisticOracleClientV2FilterDisputeablestill has an open// TODO interpret price values considering UMIPS and magic numbers, so it isn't yet interpreting proposals the way the resolution rules define.
Because of this, there's currently no way to answer a basic question: if we plugged a real strategy in, how often would it actually get the dispute decision right? A dispute bot is only trustworthy if its accuracy is measurable, and right now correctness could only be judged by manual spot-checks. A backtest harness turns that into a repeatable, numeric signal, and it becomes a regression guard, so a future change to a strategy or prompt can't quietly make the bot worse without CI catching it.
Since the package has had no functional changes since 2023, I also want to check whether it's still maintained before putting work into it.
How (proposed, pending maintainer input)
Generalize Backtest from a one-request check into a batch evaluator that:
- Replays a checked-in fixture of past Optimistic Oracle requests that have already been resolved on-chain, where the resolved value is the ground-truth "correct answer."
- Runs any pluggable
LLMDisputerStrategyover them and records what it would have decided. - Reports simple, standard quality numbers: overall accuracy, plus how often it correctly flags a bad proposal (recall) versus how often its dispute flags are actually justified (precision). In this setting a wrongful dispute and a missed bad proposal are both costly, so both numbers matter.
It would run deterministically under the existing hardhat test setup, with no live network or model calls needed for the fixture-based run, so it can gate regressions in CI. It stays strategy-agnostic: it evaluates whatever strategy is passed in and exposes no proprietary dispute logic itself, only the measurement scaffolding around a strategy.
Resolution
This is considered done when there is a merged, tested batch-backtest utility plus a small fixture of resolved requests, that scores any strategy against the known outcomes and fails CI on a quality regression. As a first step, a maintainer confirms whether contributions to llm-bot are still welcome, or whether the package is effectively deprecated in favor of internal tooling.
Difficulty Score [1-10]
5 (~4 hours) for the harness and fixture. Aligning on this issue is step 0.
cc @mrice32 sorry, not sure who to tag here.
- 主要语言
- JavaScript
- 星标
- 486
- 派生
- 215
- 平均合并
- 50 分钟
- 30 天内合并 PR
- 2
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
UMAprotocol/protocol 的其他 Issue
-
难度 5/5 一周以上 新手友好度 15/100
UMAprotocol/protocol#4931 ·
-
Add Forge Foundry Support可能重新可做 @lucifer1017 于 387 天前认领,目前没有进行中的 PR。 未关闭enhancement
难度 4/5 3-5 天 新手友好度 35/100
UMAprotocol/protocol#4802 · 1 条评论 ·
-
enhancement
难度 4/5 3-5 天 新手友好度 35/100
UMAprotocol/protocol#4794 ·
查看 UMAprotocol/protocol 的全部 Issue
相似的 Issue
-
check:passed feeds:remove
难度 1/5 1 小时以内 新手友好度 65/100
iptv-org/database#37176 · 1 条评论 · 1 个 reaction ·
维护者通常 9 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
hawk-digital-environments/HAWKI#443 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
维护者通常 1 天内回复
-
feedback simulation workshop
难度 2/5 1-3 小时 新手友好度 66/100
githubnext/gh-aw-workshop#4370 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
AltimateAI/vscode-dbt-power-user#2089 ·