Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Backtest/eval harness for the LLM dispute bot (llm-bot)

未关闭
#4,958 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
冷清
技术栈
javascript

调研方向

First confirm with a maintainer whether packages/llm-bot is still maintained. Then inspect Backtest.test, DisputerStrategy.process, and src/core/OptimisticOracleV2.ts under the existing hardhat test setup; use the resolved-request fixture and strategy interface described in the issue as the scope. Done means a tested batch utility reports accuracy, recall, and precision and CI detects regressions without live network or model calls.

由索引模型根据 Issue 内容生成。

描述

Why

The packages/llm-bot package is meant to automate a watchdog role on UMA's Optimistic Oracle: watch for proposed answers and use an LLM to decide whether a proposal looks wrong and should be disputed. Today that decision logic in the public repo is a placeholder rather than a working evaluator:

  • DisputerStrategy.process (src/core/OptimisticOracleV2.ts) doesn't call any model. It returns a fixed, hardcoded verdict every time (always "the correct answer is 1, and yes, dispute it").
  • Backtest.test can only check one request at a time against its final resolved value.
  • OptimisticOracleClientV2FilterDisputeable still has an open // TODO interpret price values considering UMIPS and magic numbers, so it isn't yet interpreting proposals the way the resolution rules define.

Because of this, there's currently no way to answer a basic question: if we plugged a real strategy in, how often would it actually get the dispute decision right? A dispute bot is only trustworthy if its accuracy is measurable, and right now correctness could only be judged by manual spot-checks. A backtest harness turns that into a repeatable, numeric signal, and it becomes a regression guard, so a future change to a strategy or prompt can't quietly make the bot worse without CI catching it.

Since the package has had no functional changes since 2023, I also want to check whether it's still maintained before putting work into it.

How (proposed, pending maintainer input)

Generalize Backtest from a one-request check into a batch evaluator that:

  1. Replays a checked-in fixture of past Optimistic Oracle requests that have already been resolved on-chain, where the resolved value is the ground-truth "correct answer."
  2. Runs any pluggable LLMDisputerStrategy over them and records what it would have decided.
  3. Reports simple, standard quality numbers: overall accuracy, plus how often it correctly flags a bad proposal (recall) versus how often its dispute flags are actually justified (precision). In this setting a wrongful dispute and a missed bad proposal are both costly, so both numbers matter.

It would run deterministically under the existing hardhat test setup, with no live network or model calls needed for the fixture-based run, so it can gate regressions in CI. It stays strategy-agnostic: it evaluates whatever strategy is passed in and exposes no proprietary dispute logic itself, only the measurement scaffolding around a strategy.

Resolution

This is considered done when there is a merged, tested batch-backtest utility plus a small fixture of resolved requests, that scores any strategy against the known outcomes and fails CI on a quality regression. As a first step, a maintainer confirms whether contributions to llm-bot are still welcome, or whether the package is effectively deprecated in favor of internal tooling.

Difficulty Score [1-10]

5 (~4 hours) for the harness and fixture. Aligning on this issue is step 0.

cc @mrice32 sorry, not sure who to tag here.

主要语言
JavaScript
星标
486
派生
215
平均合并
50 分钟
30 天内合并 PR
2

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

UMAprotocol/protocol 的其他 Issue

查看 UMAprotocol/protocol 的全部 Issue

相似的 Issue

更多 JavaScript Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。