Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Intermittent HTTP 422 on deepseek/deepseek-v4.1-flash for prompts above ~256K tokens (provider plan)

Đã đóng
#955 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Lĩnh vực
api, backend

Hướng nghiên cứu

Start at the provider gateway's /provider/v1 OpenAI-compatible chat-completions route and reproduce repeated requests above 256K tokens for deepseek/deepseek-v4.1-flash using the stated parameters. Done means oversized prompts receive a classified context error, temporary backend failures are retryable, and repeated requests never produce an unclassified 422.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

Since 2026-09-29, deepseek/deepseek-v4.1-flash on the provider plan rejects a significant share of requests with HTTP 422 and a generic body that names no field and no limit. The failures are size-correlated but not deterministic: they appear only once the prompt exceeds roughly 256K tokens, and requests of the same or larger size succeed seconds apart on the same connection settings. The same model and parameters handled 550K-608K token prompts without errors from 2026-09-25 to 2026-09-28. Impact: about 23% of requests above ~256K tokens fail (13 of 56 on 2026-09-29), and because the error looks like an invalid-request failure the client cannot tell it apart from a context overflow, so the turn is aborted and work has to be resumed manually. Long agent sessions are not reliable on this model right now.

Expected Behavior

A chat/completions request that fits the model's advertised context window returns 200 and a completion, as it did until 2026-09-28 for prompts up to 608,097 tokens. If the prompt genuinely cannot be served, the error is distinguishable: a specific status and a machine-readable cause (for example 400 with code=context_length_exceeded and the limit in the message, or 413), so a client can compress the conversation and retry instead of treating the failure as a malformed request. Requests of the same size behave consistently, regardless of which backend in the pool serves them.

Actual Behavior

The request returns HTTP 422. The body is a JSON object whose "message" field is itself a JSON-encoded string:

outer: {"message": "", "type": "server_error"}
inner (decoded): {"message":"invalid request error trace_id: <trace_id>","type":"invalid_request_error"}

The body is identical apart from the trace_id in every occurrence, and carries no field name, no limit and no limit value, so it is not possible to tell whether this is request validation or a routing/backend failure.

Observed on 2026-09-29, counting distinct requests (deduplicated by response id / trace id):

Prompt size Succeeded Failed (422)
< 262,144 tokens 1,292 0
> 262,144 tokens 43 13

Every rejection occurred in a session whose preceding successful request had already reached at least 265,955 tokens; no rejection occurred below 262,144 tokens; the largest prompt that succeeded that day was 330,799 tokens, so it is not a hard cutoff either.

Failures and successes interleave at the same size in the same session:

  • 13:47:03 rejected (preceding successful prompt ~286,800 tokens)
  • 13:47:18 succeeded, with a LARGER prompt (287,711 tokens)
  • 16:11:40 succeeded, 299,301-token prompt
  • 16:11:44 rejected, same session, four seconds later
  • 14:18:53 a summarization request was rejected
  • 14:19:16 the equivalent request succeeded, 23 seconds later

The behaviour is consistent with part of the traffic above ~256K tokens being served by a backend with a smaller context limit than the rest of the pool.

Steps to reproduce the issue
  1. Use the provider plan endpoint https://api.commandcode.ai/provider/v1/ with model deepseek/deepseek-v4.1-flash, streaming (SSE), OpenAI-compatible chat completions.
  2. Send requests with max_tokens=384000 and reasoning_effort=max, text only, no images.
  3. Let the conversation grow past ~256K prompt tokens (roughly 262,144) and keep sending single requests.
  4. Observe: requests of the same size alternate between 200 and HTTP 422 invalid request error. Below ~256K tokens the same parameters never fail. Repeating a rejected request is usually enough to get a 200 on the next attempt, which is why this is not reproducible on every attempt.

Note: max_tokens=384000 has been accepted continuously since at least 2026-09-25, including on requests that succeeded on 2026-09-29, so no parameter is statically invalid.

Command Code Version

n/a - API report, not the CLI (client details in Additional context)

Operating System

macOS

Terminal/IDE

DeepSeek Harness 0.2.0-rc.2 (web GUI), openai-completions adapter

Shell

zsh

Session file (optional)

No response

Fix prompt (optional)

Provider gateway, /provider/v1/ chat completions for OpenAI-compatible clients. When a request cannot be served - because the prompt exceeds the context limit of the backend it would be routed to, or because that backend is temporarily unusable - return a distinguishable error instead of a generic 422 invalid_request_error: 400 with type=invalid_request_error and code=context_length_exceeded plus the limit in the message when the prompt is too large, and a retryable 503/529 when the backend is temporarily unusable. Also make routing consistent: either only route a model to backends that satisfy its advertised context window, or make the failure retryable so clients can recover. Check: send the same >256K-token prompt N times and confirm every response is either a success or a classified error - never an unclassified 422 - and that repeated attempts are stable.

Additional context

Environment

  • Client: DeepSeek Harness (dsh) 0.2.0-rc.2, adapter openai-completions, streaming SSE
  • Endpoint: https://api.commandcode.ai/provider/v1/
  • Model: deepseek/deepseek-v4.1-flash (exact id as listed by your endpoint)
  • Request parameters: max_tokens=384000, reasoning_effort=max, ~35 tool definitions
  • Content: text only - the failing requests contain no images and no audio
  • Context window configured client-side: 1,000,000 tokens
  • All timestamps CEST (UTC+2)

Largest successful prompt per day (same model, same parameters)

  • 2026-09-25: 552,262 tokens
  • 2026-09-26: 550,442 tokens
  • 2026-09-27: 554,026 tokens
  • 2026-09-28: 608,097 tokens
  • 2026-09-29: 330,799 tokens, with 13 requests rejected

Change observed on 2026-09-29: the failures begin the same day that new models appeared in the catalogue served to this account (for example deepseek/deepseek-v4.1-flash-fast and claude-sonnet-5-5). No client-side change was made to this route: same model id and same request parameters as 2026-09-25 through 2026-09-28, when prompts up to 608,097 tokens completed without errors. The real maximum context window served for deepseek/deepseek-v4.1-flash now appears to be on the order of 256K rather than the 1,000,000 tokens assumed here.

Trace IDs of the rejected requests

  • 2026-09-29 01:16:58 ebeccdf9f447b92f41220911f8716881
  • 2026-09-29 13:32:45 210a01b5c40db4b23f4072bcfdc03d41
  • 2026-09-29 13:33:00 310ec7fef980d185b790161426e6dffd
  • 2026-09-29 13:35:10 50da5368dd9152f1b92f0290e6e96055
  • 2026-09-29 13:36:27 2f5b7eb22529a97a9f1c293dd5863fe7
  • 2026-09-29 13:37:07 8cbc03b6d9c337c36686a28b1fabbbe1
  • 2026-09-29 13:47:03 4da0bd59c98189521b1b7346bd4f2a5a
  • 2026-09-29 13:47:06 268d2f9e881def6bba70c7890c234e08
  • 2026-09-29 13:47:22 664e01bd3c75f7b335866e9b205f2c3d
  • 2026-09-29 13:50:36 7ec00d3a8308343f197448dc096bae51
  • 2026-09-29 14:18:57 6af735c8e2c7eb963e921961b54e8ff7
  • 2026-09-29 16:11:44 2826b95dd58064df6004d3ce82d95752
  • 2026-09-29 17:44:31 d0a9d8460d70088afa95de0264ded495

Additional request-level details for any of the trace IDs above can be provided on request.

Ngôn ngữ chính
Không có dữ liệu ngôn ngữ
Star
4k
Fork
357
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của CommandCodeAI/command-code

Tất cả issue của CommandCodeAI/command-code

Issue tương tự

Thêm issue về Backend & API Design

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.