BYOK: reasoning_content gets dropped from history, breaks multi-turn calls to thinking-mode models (400 Invalid_request_error)
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Anfängerfreundlichkeit
- 58/100
Rechercherichtung
Starte beim BYOK openai-compatible-Anfragepfad, der Bodies für /v1/chat/completions erstellt, unter Verwendung von ~/.commandcode/providers.json und der bereitgestellten Reproduktion mit mehreren Turns. Als abgeschlossen gilt die Aufgabe, wenn reasoning_content aus einer assistant-Antwort in der nächsten Anfrage neben content erhalten bleibt, sodass Thinking-Mode-Modelle den gemeldeten 400 nicht mehr zurückgeben.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Summary
Using a custom BYOK provider in ~/.commandcode/providers.json (plain openai-compatible wire type) pointed at a thinking-mode model (DeepSeek V4 Flash, GLM 5.3), a session works fine at first. At some point later in the same conversation, once the assistant's previous turn involved real generation and has to be sent back as history, it fails with:
Error: 400 Invalid_request_error The reasoning_content in the thinking mode must be passed back to the API.
What's actually happening
These models return both content and reasoning_content on the assistant message, e.g.:
{
"message": {
"role": "assistant",
"content": "pong",
"reasoning_content": "The user wants the single word "pong". Simple."
}
}
Their API requires reasoning_content to be sent back exactly as received when that message is replayed as history on a later turn. Command Code's client only seems to keep content when it rebuilds the assistant turn for the next request, so reasoning_content gets dropped and the upstream model rejects the call.
Verified this isn't the proxy
I run a small local reverse proxy in front of the provider (plain byte passthrough, forwards request bodies and streamed responses unmodified, only touches headers and drops stray literal null SSE events). To rule it out, I added a temporary log line right where it reads the incoming request body, before anything else happens to it, and drove two turns through Command Code:
[debug] /v1/chat/completions body has reasoning_content=False len=71435 (turn 1, no history yet, expected)
[debug] /v1/chat/completions body has reasoning_content=False len=71537 (turn 2, has history, still missing)
reasoning_content is already absent from the request body the moment it reaches the proxy. Since the proxy never parses or rewrites request bodies, this confirms the client drops it before the request is even sent.
One nuance: a two-turn test with short, trivial answers ("say the word alpha" / "now say beta") went through fine, no error. The failure only showed up after a turn where the model actually generated real output (in my case, a chunk of code). So the upstream seems to only enforce the reasoning_content requirement when the prior turn did substantial reasoning, not on every second turn.
Steps to reproduce
- Add a BYOK provider pointed at any OpenAI-compatible endpoint running a thinking-mode model that returns reasoning_content (DeepSeek's own API reproduces this directly, no third party needed).
- Start an interactive cmdc session on that model.
- Ask it to do something that requires actual generation (write some code, explain something at length).
- Send a follow-up in the same conversation.
- Get the 400 above.
Expected behavior
reasoning_content should be kept and sent back with content when a prior assistant turn is replayed, same as content already is.
Environment
- Command Code 1.54.2
- Windows 11 Pro (10.0.22631)
- BYOK, openai-compatible wire type
- Models: deepseek-v4-flash, glm-5.3
- Trace IDs: 634363e49a74aecf70b17642bd756acc, e362e38e22b160d3f0464663323079c1
Workaround
Switching to a model on the same provider that doesn't return reasoning_content (a Claude or GPT-class model) avoids it, since there's nothing to replay.
- Vorherrschende Sprache
- Keine Sprachdaten
- Sterne
- 4k
- Forks
- 350
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus CommandCodeAI/command-code
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#903 · 2 Kommentare ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#855 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
CommandCodeAI/command-code#841 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#655 · 1 Kommentar ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
CommandCodeAI/command-code#608 ·
Alle Issues in CommandCodeAI/command-code
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
run-llama/llama_index#23199 ·
-
p:2-high pydanty:bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 90/100
pydantic/pydantic-ai#8642 · 2 Kommentare ·
-
bug untriaged
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
opensearch-project/ml-commons#5094 ·
-
bug external groq
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
langchain-ai/langchain#40771 · 1 Kommentar ·
-
ai-observability bug team/ai-observability
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100