Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Intermittent HTTP 422 on deepseek/deepseek-v4.1-flash for prompts above ~256K tokens (provider plan)

クローズ
#955 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
領域
api, backend

調査の方向性

Start at the provider gateway's /provider/v1 OpenAI-compatible chat-completions route and reproduce repeated requests above 256K tokens for deepseek/deepseek-v4.1-flash using the stated parameters. Done means oversized prompts receive a classified context error, temporary backend failures are retryable, and repeated requests never produce an unclassified 422.

索引モデルが issue の本文から書いたものです。

説明

Summary

Since 2026-09-29, deepseek/deepseek-v4.1-flash on the provider plan rejects a significant share of requests with HTTP 422 and a generic body that names no field and no limit. The failures are size-correlated but not deterministic: they appear only once the prompt exceeds roughly 256K tokens, and requests of the same or larger size succeed seconds apart on the same connection settings. The same model and parameters handled 550K-608K token prompts without errors from 2026-09-25 to 2026-09-28. Impact: about 23% of requests above ~256K tokens fail (13 of 56 on 2026-09-29), and because the error looks like an invalid-request failure the client cannot tell it apart from a context overflow, so the turn is aborted and work has to be resumed manually. Long agent sessions are not reliable on this model right now.

Expected Behavior

A chat/completions request that fits the model's advertised context window returns 200 and a completion, as it did until 2026-09-28 for prompts up to 608,097 tokens. If the prompt genuinely cannot be served, the error is distinguishable: a specific status and a machine-readable cause (for example 400 with code=context_length_exceeded and the limit in the message, or 413), so a client can compress the conversation and retry instead of treating the failure as a malformed request. Requests of the same size behave consistently, regardless of which backend in the pool serves them.

Actual Behavior

The request returns HTTP 422. The body is a JSON object whose "message" field is itself a JSON-encoded string:

outer: {"message": "", "type": "server_error"}
inner (decoded): {"message":"invalid request error trace_id: <trace_id>","type":"invalid_request_error"}

The body is identical apart from the trace_id in every occurrence, and carries no field name, no limit and no limit value, so it is not possible to tell whether this is request validation or a routing/backend failure.

Observed on 2026-09-29, counting distinct requests (deduplicated by response id / trace id):

Prompt size Succeeded Failed (422)
< 262,144 tokens 1,292 0
> 262,144 tokens 43 13

Every rejection occurred in a session whose preceding successful request had already reached at least 265,955 tokens; no rejection occurred below 262,144 tokens; the largest prompt that succeeded that day was 330,799 tokens, so it is not a hard cutoff either.

Failures and successes interleave at the same size in the same session:

  • 13:47:03 rejected (preceding successful prompt ~286,800 tokens)
  • 13:47:18 succeeded, with a LARGER prompt (287,711 tokens)
  • 16:11:40 succeeded, 299,301-token prompt
  • 16:11:44 rejected, same session, four seconds later
  • 14:18:53 a summarization request was rejected
  • 14:19:16 the equivalent request succeeded, 23 seconds later

The behaviour is consistent with part of the traffic above ~256K tokens being served by a backend with a smaller context limit than the rest of the pool.

Steps to reproduce the issue
  1. Use the provider plan endpoint https://api.commandcode.ai/provider/v1/ with model deepseek/deepseek-v4.1-flash, streaming (SSE), OpenAI-compatible chat completions.
  2. Send requests with max_tokens=384000 and reasoning_effort=max, text only, no images.
  3. Let the conversation grow past ~256K prompt tokens (roughly 262,144) and keep sending single requests.
  4. Observe: requests of the same size alternate between 200 and HTTP 422 invalid request error. Below ~256K tokens the same parameters never fail. Repeating a rejected request is usually enough to get a 200 on the next attempt, which is why this is not reproducible on every attempt.

Note: max_tokens=384000 has been accepted continuously since at least 2026-09-25, including on requests that succeeded on 2026-09-29, so no parameter is statically invalid.

Command Code Version

n/a - API report, not the CLI (client details in Additional context)

Operating System

macOS

Terminal/IDE

DeepSeek Harness 0.2.0-rc.2 (web GUI), openai-completions adapter

Shell

zsh

Session file (optional)

No response

Fix prompt (optional)

Provider gateway, /provider/v1/ chat completions for OpenAI-compatible clients. When a request cannot be served - because the prompt exceeds the context limit of the backend it would be routed to, or because that backend is temporarily unusable - return a distinguishable error instead of a generic 422 invalid_request_error: 400 with type=invalid_request_error and code=context_length_exceeded plus the limit in the message when the prompt is too large, and a retryable 503/529 when the backend is temporarily unusable. Also make routing consistent: either only route a model to backends that satisfy its advertised context window, or make the failure retryable so clients can recover. Check: send the same >256K-token prompt N times and confirm every response is either a success or a classified error - never an unclassified 422 - and that repeated attempts are stable.

Additional context

Environment

  • Client: DeepSeek Harness (dsh) 0.2.0-rc.2, adapter openai-completions, streaming SSE
  • Endpoint: https://api.commandcode.ai/provider/v1/
  • Model: deepseek/deepseek-v4.1-flash (exact id as listed by your endpoint)
  • Request parameters: max_tokens=384000, reasoning_effort=max, ~35 tool definitions
  • Content: text only - the failing requests contain no images and no audio
  • Context window configured client-side: 1,000,000 tokens
  • All timestamps CEST (UTC+2)

Largest successful prompt per day (same model, same parameters)

  • 2026-09-25: 552,262 tokens
  • 2026-09-26: 550,442 tokens
  • 2026-09-27: 554,026 tokens
  • 2026-09-28: 608,097 tokens
  • 2026-09-29: 330,799 tokens, with 13 requests rejected

Change observed on 2026-09-29: the failures begin the same day that new models appeared in the catalogue served to this account (for example deepseek/deepseek-v4.1-flash-fast and claude-sonnet-5-5). No client-side change was made to this route: same model id and same request parameters as 2026-09-25 through 2026-09-28, when prompts up to 608,097 tokens completed without errors. The real maximum context window served for deepseek/deepseek-v4.1-flash now appears to be on the order of 256K rather than the 1,000,000 tokens assumed here.

Trace IDs of the rejected requests

  • 2026-09-29 01:16:58 ebeccdf9f447b92f41220911f8716881
  • 2026-09-29 13:32:45 210a01b5c40db4b23f4072bcfdc03d41
  • 2026-09-29 13:33:00 310ec7fef980d185b790161426e6dffd
  • 2026-09-29 13:35:10 50da5368dd9152f1b92f0290e6e96055
  • 2026-09-29 13:36:27 2f5b7eb22529a97a9f1c293dd5863fe7
  • 2026-09-29 13:37:07 8cbc03b6d9c337c36686a28b1fabbbe1
  • 2026-09-29 13:47:03 4da0bd59c98189521b1b7346bd4f2a5a
  • 2026-09-29 13:47:06 268d2f9e881def6bba70c7890c234e08
  • 2026-09-29 13:47:22 664e01bd3c75f7b335866e9b205f2c3d
  • 2026-09-29 13:50:36 7ec00d3a8308343f197448dc096bae51
  • 2026-09-29 14:18:57 6af735c8e2c7eb963e921961b54e8ff7
  • 2026-09-29 16:11:44 2826b95dd58064df6004d3ce82d95752
  • 2026-09-29 17:44:31 d0a9d8460d70088afa95de0264ded495

Additional request-level details for any of the trace IDs above can be provided on request.

主要言語
言語のデータがありません
スター
4k
フォーク
357
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

CommandCodeAI/command-code のほかの issue

CommandCodeAI/command-code の issue をすべて見る

似ている issue

Backend & API Design の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。