Backtest/eval harness for the LLM dispute bot (llm-bot)
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- javascript
- Área
- blockchain, testing-qa
Línea de trabajo
First confirm with a maintainer whether packages/llm-bot is still maintained. Then inspect Backtest.test, DisputerStrategy.process, and src/core/OptimisticOracleV2.ts under the existing hardhat test setup; use the resolved-request fixture and strategy interface described in the issue as the scope. Done means a tested batch utility reports accuracy, recall, and precision and CI detects regressions without live network or model calls.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Why
The packages/llm-bot package is meant to automate a watchdog role on UMA's Optimistic Oracle: watch for proposed answers and use an LLM to decide whether a proposal looks wrong and should be disputed. Today that decision logic in the public repo is a placeholder rather than a working evaluator:
DisputerStrategy.process(src/core/OptimisticOracleV2.ts) doesn't call any model. It returns a fixed, hardcoded verdict every time (always "the correct answer is 1, and yes, dispute it").Backtest.testcan only check one request at a time against its final resolved value.OptimisticOracleClientV2FilterDisputeablestill has an open// TODO interpret price values considering UMIPS and magic numbers, so it isn't yet interpreting proposals the way the resolution rules define.
Because of this, there's currently no way to answer a basic question: if we plugged a real strategy in, how often would it actually get the dispute decision right? A dispute bot is only trustworthy if its accuracy is measurable, and right now correctness could only be judged by manual spot-checks. A backtest harness turns that into a repeatable, numeric signal, and it becomes a regression guard, so a future change to a strategy or prompt can't quietly make the bot worse without CI catching it.
Since the package has had no functional changes since 2023, I also want to check whether it's still maintained before putting work into it.
How (proposed, pending maintainer input)
Generalize Backtest from a one-request check into a batch evaluator that:
- Replays a checked-in fixture of past Optimistic Oracle requests that have already been resolved on-chain, where the resolved value is the ground-truth "correct answer."
- Runs any pluggable
LLMDisputerStrategyover them and records what it would have decided. - Reports simple, standard quality numbers: overall accuracy, plus how often it correctly flags a bad proposal (recall) versus how often its dispute flags are actually justified (precision). In this setting a wrongful dispute and a missed bad proposal are both costly, so both numbers matter.
It would run deterministically under the existing hardhat test setup, with no live network or model calls needed for the fixture-based run, so it can gate regressions in CI. It stays strategy-agnostic: it evaluates whatever strategy is passed in and exposes no proprietary dispute logic itself, only the measurement scaffolding around a strategy.
Resolution
This is considered done when there is a merged, tested batch-backtest utility plus a small fixture of resolved requests, that scores any strategy against the known outcomes and fails CI on a quality regression. As a first step, a maintainer confirms whether contributions to llm-bot are still welcome, or whether the package is effectively deprecated in favor of internal tooling.
Difficulty Score [1-10]
5 (~4 hours) for the harness and fixture. Aligning on this issue is step 0.
cc @mrice32 sorry, not sure who to tag here.
- Lenguaje dominante
- JavaScript
- Estrellas
- 486
- Forks
- 215
- Merge medio
- 50 min
- PR fusionados (30 d)
- 2
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de UMAprotocol/protocol
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
UMAprotocol/protocol#4931 ·
-
Add Forge Foundry SupportQuizá libre de nuevo @lucifer1017 la tomó hace 387 días y no hay ningún pull request abierto. Abiertoenhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
UMAprotocol/protocol#4802 · 1 comentario ·
-
enhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
UMAprotocol/protocol#4794 ·
Todos los issues de UMAprotocol/protocol
Issues similares
-
Remove: Fox Deportes SDAbiertocheck:passed feeds:remove
Dificultad 1/5 Menos de una hora Aptitud para principiantes 65/100
iptv-org/database#37176 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 9 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
hawk-digital-environments/HAWKI#443 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Los mantenedores suelen responder en 1 día
-
feedback simulation workshop
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
githubnext/gh-aw-workshop#4370 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
AltimateAI/vscode-dbt-power-user#2089 ·