Why Do Results from MCP-Trained Models Differ Greatly Between generate_benchmarks.py and train.py
维护者通常 3 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 32/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
- 技术栈
- python
调研方向
首先比较 generate_benchmarks.py 和 train.py,重点关注它们各自如何处理 complete_task 工具以及如何确定任务是否成功。跟踪经过 MCP 训练的模型的评估路径,并使用提供的 vLLM 和 LoRA 配置复现这一差异。完成的标准是确认完成条件是否存在差异,并记录或修复验证缺口的原因。
由索引模型根据 Issue 内容生成。
描述
Why does a model trained via MCP show a significant discrepancy in results when tested using generate_benchmarks.py, compared to the outcomes from train.py?
A preliminary investigation indicates that the root cause may be related to the following logic: In generate_benchmarks.py, the model must call the complete_task function to be deemed as having finished a task. However, there is no such logic implemented in train.py. Is this the reason for the large result deviation?
generate_benchmarks.py
qwen3_4b_instruct = art.Model(
name="qwen3-4b-instruct",
project=server,
inference_model_name="qwen3-4b-instruct",
inference_base_url="http://localhost:8082/v1", #http://localhost:8082/v1
inference_api_key="dummy", # API key
inference_timeout=3600,
)
source /*****/miniconda3/bin/activate ART
BASE_MODEL_PATH="/*****/Qwen3-4B-Instruct-2507"
LORA_PATH="/*****/examples/mcp-rl/.art/mcp-agent-training/models/mcp-4b-001/checkpoints/0017"
CUDA_VISIBLE_DEVICES=2,3 python -m vllm.entrypoints.openai.api_server \
--model "$BASE_MODEL_PATH" \
--served-model-name "mcp-4b-001-finetuned" \
--enable-lora \
--lora-modules mcp-4b-001="$LORA_PATH" \
--host 0.0.0.0 \
--port 8082 \
--trust-remote-code \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--max-model-len 16384 \
--tensor-parallel-size 2
In the generate_benchmarks.py
mcp-4b-001-finetuned = art.Model(
name="mcp-4b-001-finetuned",
project=server,
inference_model_name="mcp-4b-001-finetuned",
inference_base_url="http://localhost:8082/v1", #http://localhost:8082/v1
inference_api_key="dummy",
inference_timeout=3600)
While this method has a certain degree of effectiveness, there is still a significant gap between its current performance and the validation results obtained during training.
I would like to know: during the training process, is it also mandatory for the model to output the "complete task" tool to be considered a successful completion of the task? Because when I used your benchmark, the trained large model tended not to call the "complete task" tool to end the task, resulting in an evaluation success rate of 0.
- 主要语言
- Python
- 星标
- 10.8k
- 派生
- 989
- 平均合并
- 11 小时 38 分钟
- 30 天内合并 PR
- 104
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
OpenPipe/ART 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 3 天内回复
-
难度 4/5 3-5 天 新手友好度 54/100
维护者通常 3 天内回复
-
难度 5/5 一周以上 新手友好度 25/100
维护者通常 3 天内回复
-
难度 5/5 一周以上 新手友好度 10/100
维护者通常 3 天内回复
-
难度 5/5 一周以上 新手友好度 42/100
维护者通常 3 天内回复
相似的 Issue
-
docs pydanty:is-working
难度 2/5 1-3 小时 新手友好度 75/100
pydantic/pydantic-ai#8863 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
run-llama/llama_index#23278 ·
维护者通常 2 天内回复
-
documentation from-review-extraction github-actions priority: low severity:nit
难度 1/5 1 小时以内 新手友好度 92/100
LearningCircuit/local-deep-research#6946 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
oracle/langchain-oracle#323 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 88/100
tenstorrent/tt-metal#58057 · 1 条评论 ·
维护者通常 1 天内回复