MiniMax-M2.7 / M3: tool output not propagated between chained tool calls (design_workflow → validate_design)
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Área
- ai
Línea de trabajo
Comienza con la configuración de harness-test-bench y el escenario smoke_all_tools_2; después inspecciona la Task 2 (task_2_module_design) y sus llamadas design_workflow → validate_design. Compara MiniMax-M2.7/M3 con M2.5; se considera terminado cuando la llamada encadenada recibe el resultado exacto de design_workflow en una solicitud validate_design válida según el esquema.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Capability area: Agent harness / Tool use
What does M2.7 fail to do for you?
MiniMax-M2.7 fails to produce a valid JSON schema conforming to the validate_design tool definition when chaining from design_workflow.
What would "good" look like in M3?
validate_design receives a valid JSON schema that conforms to the tool's parameter definition — populated from the actual design_workflow result, not reconstructed from model reasoning.
Summary
MiniMax-M2.7 and MiniMax-M3 both fail to propagate tool output across a sequential tool call chain. When design_workflow returns a schema object, both models call the next tool (validate_design) with a schema that does not match the returned output — either reconstructed, incomplete, or structurally invalid against the tool definition.
MiniMax-M2.5 does not reproduce this behavior.
Environment
| Field | Value |
|---|---|
| Models affected | MiniMax-M2.7, MiniMax-M3 |
| Model passing | MiniMax-M2.5 |
| Endpoint | OpenRouter |
| Tool calling | Sequential 2-step chain |
| Observed | 2026-06-17 (UTC+8) |
Expected behavior
design_workflowis called and returns a schema object- The model reads the tool result
validate_designis called withschemaset to the exact object returned in step 1, conforming to the tool's JSON schema definition
Actual behavior
The model calls validate_design with a schema argument that does not match the design_workflow output. The model appears to reconstruct the schema from internal reasoning rather than reading the tool result, producing a value that violates the tool's parameter schema — missing required fields or structurally incorrect.
Benchmark eval output:
issues: ["Validated design is missing fields: project name, client, start date, deadline, owner, budget"]
Tool call sequence observed
1. design_workflow(...)
→ returns: { name, information: [{name, type, ai_hint}, ...], states: [...], ... }
2. validate_design({ schema: <does not match step 1 output, fails tool schema validation> })
→ missing required information fields
Steps to reproduce
We have an open-source MCP tool benchmark that reproduces this consistently:
Repo: https://github.com/Inistate/harness-test-bench
- Clone the repo and follow the setup instructions
- Run the
smoke_all_tools_2scenario againstminimax/MiniMax-M2.7orminimax/MiniMax-M3 - Inspect Task 2 (
task_2_module_design) —validate_designwill be called with a schema that does not conform to the tool's parameter definition - Compare against a passing model (e.g.
minimax/MiniMax-M2.5) running the same scenario
Questions
- Is this a known regression from M2.5?
- Does the model read tool results back into context between sequential calls, or is this not guaranteed?
- Is there a recommended pattern to ensure tool output is forwarded as-is rather than reconstructed?
Workaround
We currently exclude M2.7 and M3 from workflows requiring sequential tool output propagation. A fix at the model level would be preferable.
References
Reproduced across multiple runs in an independent MCP tool benchmark. M2.5 passed consistently. M2.7 scored 3/5 and M3 scored 2/5, with the T2 chain failing on every run for both models.
- Lenguaje dominante
- Sin datos de lenguaje
- Estrellas
- 489
- Forks
- 59
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de MiniMax-AI/MiniMax-M3
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
MiniMax-AI/MiniMax-M3#34 · 1 reacción ·
-
[M2.7 Bug]Abiertobug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
MiniMax-AI/MiniMax-M3#35 ·
-
enhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
MiniMax-AI/MiniMax-M3#33 ·
-
Verify evals on Papers with CodeAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 48/100
MiniMax-AI/MiniMax-M3#32 ·
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 42/100
MiniMax-AI/MiniMax-M3#31 ·
Todos los issues de MiniMax-AI/MiniMax-M3
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
higress-group/higress#4839 ·
Los mantenedores suelen responder en 1 día
-
org:external package:deepagents priority:backlog topic:models topic:multimodal type:feature
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
langchain-ai/deepagents#6599 ·
Los mantenedores suelen responder en 1 día
-
p:3-mid pydanty:pr
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pydantic/pydantic-ai#8968 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-2 días Aptitud para principiantes 84/100
bytedance/trae-agent#504 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
0xPlaygrounds/rig#2621 · 1 comentario ·
Los mantenedores suelen responder en 1 día