Anthropic API: tool calls intermittently emitted inside `thinking` and returned as `end_turn`
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
Línea de trabajo
Start with the attached failing SSE log and trace the /v1/messages streaming path for thinking and tool-call output. Compare a failing turn with thinking enabled against a successful tool-use turn and the reported thinking-off run. Done means intended tool calls are emitted in an Anthropic tool_use block with stop_reason: "tool_use", rather than as thinking_delta followed by end_turn.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
AI disclosure: This issue was created with the assistance of AI based on captured Pi session logs and raw SSE traffic. Im not a dev :)
System
- OS: Ubuntu 26.04.1 LTS
- Kernel: 7.0.0-34-generic
- CPU: AMD Ryzen 7 9800X3D
- RAM: 32 GB DDR5-6000
- GPU: NVIDIA GeForce RTX 3090 24 GB
- NVIDIA driver: 595.91.07
- CUDA: 13.2
- Strata: Docker, v0.1.38
- Model:
qwen3.8-flash-next-coder-iq1_m - Context: 262144
- API: Anthropic-compatible
/v1/messages - Client: Pi 1.0.2
- Pi thinking mode:
low(budget_tokens: 2048)
Setup
- Strata: current Docker build / v0.1.38-era code
- Model:
qwen3.8-flash-next-coder-iq1_m - API:
/v1/messages - Client: Pi 1.0.2
- Pi thinking:
low - Request thinking budget:
2048 - Context: well below 262K
With thinking enabled, Strata intermittently emits a complete intended tool call inside the Anthropic thinking block instead of producing a structured tool_use.
Raw SSE from a failing request:
event: content_block_delta
data: {"type":"content_block_delta","index":0,
"delta":{"type":"thinking_delta","thinking":"<tool_call>"}}
...
event: content_block_delta
data: {"type":"content_block_delta","index":0,
"delta":{"type":"thinking_delta","thinking":"</tool_call>"}}
Immediately afterwards:
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta",
"delta":{"stop_reason":"end_turn","stop_sequence":null},
"usage":{"input_tokens":243,
"cache_read_input_tokens":28792,
"output_tokens":79}}
event: message_stop
data: {"type":"message_stop"}
There is no Anthropic tool_use content block in the failed turn.
Successful turns in the same session correctly produce:
event: content_block_start
data: {"type":"content_block_start","index":2,
"content_block":{"type":"tool_use", ...}}
event: message_delta
data: {"type":"message_delta",
"delta":{"stop_reason":"tool_use", ...}}
Reproduction
Using the same multi-step tool workload:
thinking low: C1 FAIL / C2 FAIL / C3 FAIL
thinking off: C4 PASS
Failures occurred at different context sizes (~21K–34K), so this does not appear to be a context-limit or timeout issue.
Expected
When the model transitions from reasoning to a tool call, Strata should emit a structured Anthropic tool_use block and stop_reason: "tool_use".
Actual
The intended tool call is sometimes streamed as thinking_delta, followed by a clean end_turn, causing the client agent loop to stop.
- Lenguaje dominante
- C++
- Estrellas
- 11.6k
- Forks
- 1k
- Merge medio
- 7 h 46 min
- PR fusionados (30 d)
- 30
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Posiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 83/100
Los mantenedores suelen responder en 1 día
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Posiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
Los mantenedores suelen responder en 1 día
-
MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
Todos los issues de Niko1221/Strata
Issues similares
-
agent:Windows bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
? - Needs Triage bot_watch bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
NVIDIA/cudf-spark-jni#5267 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día
-
(S1 - Need confirmation)
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
CleverRaven/Cataclysm-DDA#88974 ·
Los mantenedores suelen responder en 1 día