Send us a transcript where backcheck got it wrong
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 72/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- rust
- 領域
- cli, testing-qa
調査の方向性
src/claims.rs の関連するパターン、特に is_hedged() と is_negated() から始め、次に tests/fixtures/ に最小限のサニタイズ済み JSONL fixture を用意して誤った判定を再現します。backcheck -f your-fixture.jsonl --json を実行し、報告された主張が想定どおりに分類され、過剰マッチングを引き起こさないことを示す回帰テストを追加します。
索引モデルが issue の本文から書いたものです。
説明
backcheck is only worth installing if you trust its verdicts. A hook that fires on honest work
gets uninstalled within a day, and then it protects nobody.
So the most valuable thing you can contribute is a case where it was wrong.
Two kinds of wrong
False positive — it flagged work that was fine. This is the expensive kind. Examples we
already fixed this way:
- runners invoked through a virtualenv path (
.venv/bin/python -m pytest) were invisible, so
genuine runs looked like no run at all - the shell builtin
test -fwas counted as a test run, which could hide a missing suite - "Ruff passes with no warnings" was read as a negated sentence and skipped
False negative — an agent claimed something it had not done and backcheck stayed quiet.
Usually an unrecognised runner (#3) or a claim phrasing the patterns miss.
How to report one
Please do not attach a raw transcript. They contain your source, your paths, and sometimes
your credentials.
Send the smallest JSONL that reproduces it, with everything sensitive replaced. Three lines is
usually enough, and it can go straight into tests/fixtures/ as a regression test:
{"type":"assistant","message":{"content":[{"type":"tool_use","id":"t1","name":"Bash","input":{"command":"<command>"}}]}}
{"type":"user","toolUseResult":{"stdout":"<output>","stderr":"","interrupted":false},"message":{"content":[{"type":"tool_result","tool_use_id":"t1","content":"<output>"}]}}
{"type":"assistant","message":{"content":[{"type":"text","text":"<what the agent claimed>"}]}}
Then run backcheck -f your-fixture.jsonl --json and paste the output along with what you
expected instead.
There is an issue template for this: 🎯 Wrong verdict.
Claim phrasings we know are missed
Patterns live in src/claims.rs. Known gaps:
- non-English summaries
- emoji-only status (
✅ testswith no verb) - markdown tables reporting per-check status
- "everything is green", "CI is happy", "all clear"
- claims about coverage thresholds
- claims split across sentences ("Ran the suite. Everything passed.")
Each of those is a pattern plus a test, and the guard against over-matching is the interesting
part — see is_hedged() and is_negated().
- 主要言語
- Rust
- スター
- 1
- フォーク
- 2
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
VectorInstitute/backcheck のほかの issue
-
Add support for more test runners and linters対応中かも @OllieinCanada が 55 日前に担当しました。 オープンgood first issue help wanted runner
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
-
accuracy enhancement help wanted
難易度 5/5 1週間以上 初心者へのやさしさ 32/100
VectorInstitute/backcheck#11 ·
-
enhancement good first issue help wanted
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
VectorInstitute/backcheck#10 ·
-
enhancement good first issue help wanted
難易度 3/5 1〜2日 初心者へのやさしさ 68/100
-
accuracy enhancement help wanted
難易度 4/5 3〜5日 初心者へのやさしさ 50/100
VectorInstitute/backcheck の issue をすべて見る
似ている issue
-
難易度 1/5 1〜3時間 初心者へのやさしさ 72/100
MattA-Official/vwmcp#20 ·
-
documentation good first issue
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
undergroundrap/hatchling#15 ·
-
Progress difficulty filter lists Hard before Medium対応中かも @Pandamachi が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
sysprog21/codetrial#281 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
CLI: --verdict silently ignores extra program/file arguments対応中かも @oxura が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
HigherOrderCO/Bend#1503 ·
-
Add MySQL test coverage for numeric_precision/numeric_scale and seq_in_fk (follow-up to #939)対応中かも @tosinxt が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
TabularisDB/tabularis#977 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信