Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Backtest/eval harness for the LLM dispute bot (llm-bot)

Abierto
#4,958 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
48/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
javascript

Línea de trabajo

First confirm with a maintainer whether packages/llm-bot is still maintained. Then inspect Backtest.test, DisputerStrategy.process, and src/core/OptimisticOracleV2.ts under the existing hardhat test setup; use the resolved-request fixture and strategy interface described in the issue as the scope. Done means a tested batch utility reports accuracy, recall, and precision and CI detects regressions without live network or model calls.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Why

The packages/llm-bot package is meant to automate a watchdog role on UMA's Optimistic Oracle: watch for proposed answers and use an LLM to decide whether a proposal looks wrong and should be disputed. Today that decision logic in the public repo is a placeholder rather than a working evaluator:

  • DisputerStrategy.process (src/core/OptimisticOracleV2.ts) doesn't call any model. It returns a fixed, hardcoded verdict every time (always "the correct answer is 1, and yes, dispute it").
  • Backtest.test can only check one request at a time against its final resolved value.
  • OptimisticOracleClientV2FilterDisputeable still has an open // TODO interpret price values considering UMIPS and magic numbers, so it isn't yet interpreting proposals the way the resolution rules define.

Because of this, there's currently no way to answer a basic question: if we plugged a real strategy in, how often would it actually get the dispute decision right? A dispute bot is only trustworthy if its accuracy is measurable, and right now correctness could only be judged by manual spot-checks. A backtest harness turns that into a repeatable, numeric signal, and it becomes a regression guard, so a future change to a strategy or prompt can't quietly make the bot worse without CI catching it.

Since the package has had no functional changes since 2023, I also want to check whether it's still maintained before putting work into it.

How (proposed, pending maintainer input)

Generalize Backtest from a one-request check into a batch evaluator that:

  1. Replays a checked-in fixture of past Optimistic Oracle requests that have already been resolved on-chain, where the resolved value is the ground-truth "correct answer."
  2. Runs any pluggable LLMDisputerStrategy over them and records what it would have decided.
  3. Reports simple, standard quality numbers: overall accuracy, plus how often it correctly flags a bad proposal (recall) versus how often its dispute flags are actually justified (precision). In this setting a wrongful dispute and a missed bad proposal are both costly, so both numbers matter.

It would run deterministically under the existing hardhat test setup, with no live network or model calls needed for the fixture-based run, so it can gate regressions in CI. It stays strategy-agnostic: it evaluates whatever strategy is passed in and exposes no proprietary dispute logic itself, only the measurement scaffolding around a strategy.

Resolution

This is considered done when there is a merged, tested batch-backtest utility plus a small fixture of resolved requests, that scores any strategy against the known outcomes and fails CI on a quality regression. As a first step, a maintainer confirms whether contributions to llm-bot are still welcome, or whether the package is effectively deprecated in favor of internal tooling.

Difficulty Score [1-10]

5 (~4 hours) for the harness and fixture. Aligning on this issue is step 0.

cc @mrice32 sorry, not sure who to tag here.

Lenguaje dominante
JavaScript
Estrellas
486
Forks
215
Merge medio
50 min
PR fusionados (30 d)
2

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de UMAprotocol/protocol

Todos los issues de UMAprotocol/protocol

Issues similares

Más issues de JavaScript

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.