Benchmark and improve tool-search ranking with indexed BM25
維護者通常 1 天內回覆
評估
- 難度
- 5/5
- 預估耗時
- 一週以上
- 新手友好度
- 45/100
- Issue 類型
- 功能
- 描述清晰度
- 基本清楚
- 活躍度
- 活躍
- 技術堆疊
- go
研究方向
首先定位伺服器現有的 tool-search 實作,並檢視用於索引工具名稱、描述、參數名稱和參數描述的原型 benchmark 及其單元測試。與維護者確認首選範圍,然後針對包含 49 個查詢和 115 個工具的 benchmark 測量所選的排名策略,並回報檢索品質和延遲,且不得出現回歸。
由索引模型根據 Issue 內容生成。
描述
Describe the feature or problem you’d like to solve
The GitHub MCP Server already exposes tool discovery/search functionality, but
there is no repeatable benchmark for measuring how reliably natural-language
queries retrieve the intended MCP tool.
As the tool inventory grows, a benchmark would make ranking changes measurable
and help prevent retrieval regressions.
This is separate from host-side deferred tool loading discussed in #1680. The
proposal only concerns ranking inside the server's existing tool-search
implementation.
Proposed solution
Add a hand-labelled benchmark covering natural-language intents across the
server's major toolsets, then compare the current heuristic with an indexed
BM25 implementation.
A prototype benchmark contains 49 queries over 115 unique tools and produced:
| Strategy | Recall@1 | Recall@3 | MRR@10 | Query latency |
|---|---|---|---|---|
| Current heuristic | 71.4% | 81.6% | 0.792 | ~2.25 ms |
| Indexed BM25 | 71.4% | 87.8% | 0.802 | ~34 µs |
| Hybrid RRF | 73.5% | 87.8% | 0.823 | ~2.38 ms |
Indexed BM25 improved Recall@3 by 6.1 percentage points and was approximately
66x faster per query. The hybrid produced the strongest ranking quality.
Before submitting a PR, I would appreciate maintainer guidance on the preferred
scope:
- Benchmark harness only
- Benchmark plus indexed BM25
- Benchmark plus a hybrid ranking experiment
Example prompts or workflows
- "Find open issues assigned to me across repositories"
- "Read the files, reviews, and diff for a pull request"
- "Download logs for a failed workflow job"
- "Find exposed secrets detected in a repository"
- "Add an issue to a GitHub project"
Additional context
The benchmark uses the complete current tool inventory and validates that every
labelled relevant tool exists. The prototype includes unit tests for indexing
tool names, descriptions, parameter names, and parameter descriptions.
- 主要語言
- Go
- 星號
- 33.3k
- 分支
- 5.1k
- 平均合併
- 22 小時 46 分鐘
- 30 天內合併 PR
- 17
環境準備
- 提供 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
github/github-mcp-server 的其他 Issue
-
pull_request_read drops merge_commit_sha可能已有人在做 @thejdubb02 於 27 天前認領。 未關閉bug
難度 2/5 1-3 小時 新手友好度 84/100
github/github-mcp-server#3235 · 1 則留言 ·
維護者通常 1 天內回覆
-
Add guidance on GitHub autolinked reference formatting for AI agents可能重新可做 關聯的 PR 已關閉且未合併。 未關閉enhancement
難度 1/5 1 小時以內 新手友好度 88/100
github/github-mcp-server#3042 · 2 則留言 ·
維護者通常 1 天內回覆
-
Incorrect install documentation leads to error "error: unknown option '-e'"可能已有人在做 @syf2211 於 56 天前認領。 未關閉bug
難度 2/5 1-3 小時 新手友好度 72/100
github/github-mcp-server#3032 · 1 則留言 · 1 個 reaction ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 74/100
github/github-mcp-server#2803 · 1 則留言 ·
維護者通常 1 天內回覆
-
get_discussion and get_discussion_comments accept calls with missing required parameters instead of returning a validation error可能已有人在做 @rodboev 於 106 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 76/100
github/github-mcp-server#2740 ·
維護者通常 1 天內回覆
查看 github/github-mcp-server 的全部 Issue
相似的 Issue
-
難度 1/5 1 小時以內 新手友好度 92/100
JuliusBrussee/caveman#1189 ·
維護者通常 1 天內回覆
-
agent-review-finding chore
難度 2/5 半天 新手友好度 78/100
jordansmall/spindrift#4497 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 76/100
維護者通常 1 天內回覆
-
bug from-studio
難度 2/5 1-3 小時 新手友好度 70/100
esengine/DeepSeek-Reasonix#12044 · 2 則留言 ·
維護者通常 1 天內回覆
-
難度 1/5 1-3 小時 新手友好度 82/100