Send us a transcript where backcheck got it wrong
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 72/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- rust
- Área
- cli, testing-qa
Línea de trabajo
Comienza con el patrón relevante en src/claims.rs, especialmente is_hedged() y is_negated(), y luego reproduce el veredicto falso con un fixture JSONL mínimo y saneado en tests/fixtures/. Ejecuta backcheck -f your-fixture.jsonl --json y añade una prueba de regresión que muestre que la afirmación reportada se clasifica como se espera sin provocar coincidencias excesivas.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
backcheck is only worth installing if you trust its verdicts. A hook that fires on honest work
gets uninstalled within a day, and then it protects nobody.
So the most valuable thing you can contribute is a case where it was wrong.
Two kinds of wrong
False positive — it flagged work that was fine. This is the expensive kind. Examples we
already fixed this way:
- runners invoked through a virtualenv path (
.venv/bin/python -m pytest) were invisible, so
genuine runs looked like no run at all - the shell builtin
test -fwas counted as a test run, which could hide a missing suite - "Ruff passes with no warnings" was read as a negated sentence and skipped
False negative — an agent claimed something it had not done and backcheck stayed quiet.
Usually an unrecognised runner (#3) or a claim phrasing the patterns miss.
How to report one
Please do not attach a raw transcript. They contain your source, your paths, and sometimes
your credentials.
Send the smallest JSONL that reproduces it, with everything sensitive replaced. Three lines is
usually enough, and it can go straight into tests/fixtures/ as a regression test:
{"type":"assistant","message":{"content":[{"type":"tool_use","id":"t1","name":"Bash","input":{"command":"<command>"}}]}}
{"type":"user","toolUseResult":{"stdout":"<output>","stderr":"","interrupted":false},"message":{"content":[{"type":"tool_result","tool_use_id":"t1","content":"<output>"}]}}
{"type":"assistant","message":{"content":[{"type":"text","text":"<what the agent claimed>"}]}}
Then run backcheck -f your-fixture.jsonl --json and paste the output along with what you
expected instead.
There is an issue template for this: 🎯 Wrong verdict.
Claim phrasings we know are missed
Patterns live in src/claims.rs. Known gaps:
- non-English summaries
- emoji-only status (
✅ testswith no verb) - markdown tables reporting per-check status
- "everything is green", "CI is happy", "all clear"
- claims about coverage thresholds
- claims split across sentences ("Ran the suite. Everything passed.")
Each of those is a pattern plus a test, and the guard against over-matching is the interesting
part — see is_hedged() and is_negated().
- Lenguaje dominante
- Rust
- Estrellas
- 1
- Forks
- 2
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de VectorInstitute/backcheck
-
Add support for more test runners and lintersPosiblemente ocupada @OllieinCanada la tomó hace 55 días. Abiertogood first issue help wanted runner
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
-
accuracy enhancement help wanted
Dificultad 5/5 Más de una semana Aptitud para principiantes 32/100
VectorInstitute/backcheck#11 ·
-
enhancement good first issue help wanted
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
VectorInstitute/backcheck#10 ·
-
enhancement good first issue help wanted
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
-
accuracy enhancement help wanted
Dificultad 4/5 3-5 días Aptitud para principiantes 50/100
Todos los issues de VectorInstitute/backcheck
Issues similares
-
Progress difficulty filter lists Hard before MediumPosiblemente ocupada @Pandamachi la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
sysprog21/codetrial#281 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
C-bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
rust-lang/rust-analyzer#23501 ·
Los mantenedores suelen responder en 1 día
-
Streamable HTTP client: a 401 or 403 with a JSON-RPC error body and no WWW-Authenticate loses its HTTP statusPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abiertobug P2 ready for work T-security T-transport
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
modelcontextprotocol/rust-sdk#1339 ·
Los mantenedores suelen responder en 3 días
-
scripts/gen-gallery.py:118: a ready session now reports in_progress, so SESSION_READY_OLD can goAbiertonightly-audit
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
antithesishq/snouty#396 ·
Los mantenedores suelen responder en 1 día
-
French BIP39 wordlist starts with a UTF-8 BOM, so generated French mnemonics carry U+FEFF and derive a non-canonical seedPosiblemente ocupada @Kshot3000 la tomó hoy. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 91/100
ergoplatform/sigma-rust#976 ·
Los mantenedores suelen responder en 1 día