Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: guided response_format is silently ignored without guided_decoding_backend

オープン 初心者向け
#19,635 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
82/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
api, backend

調査の方向性

BaseLLM._check_arguments() から始め、tests/unittest/llmapi/test_sampling_params.py にある関連ヘルパーを読みます。まず指定された回帰テストを実行し、次に報告されたコントロールケースを使って、guided_decoding_backend のないアクティブな guided decoding が拒否される一方で、no-op のリクエストと設定済みバックエンドを使用するリクエストでは動作が維持されることを確認します。

索引モデルが issue の本文から書いたものです。

説明

bug Inference runtime
System Info

TensorRT-LLM commit:
c76f4a856447820aa1f994b053da7580d69b657d

Container:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232

Container environment:

  • NVIDIA Release 26.08
  • PyTorch 2.14.0a0+4fdf77b
  • Python 3.12

Host:

  • Linux x86_64
  • Docker 29.1.3
  • CPU-only host for the regression test

The container reports the same TRT_LLM_GIT_COMMIT as the source revision above.

Who can help?

No response

Information
  • The official example scripts
  • My own modified scripts
Tasks
  • An officially supported task in the examples folder (such as GLUE/SQuAD, ...)
  • My own task or dataset (give details below)
Reproduction

I ran into the remaining missing-backend case from NVIDIA/TensorRT#4855 and reproduced it against current TensorRT-LLM main:

c76f4a856447820aa1f994b053da7580d69b657d

I reduced the issue to the request-vs-engine validation path in BaseLLM._check_arguments().

The regression test I used is:

def test_check_arguments_rejects_guided_decoding_without_backend() -> None:
    llm = _TestLLM(guided_decoding_backend=None)
    sampling_params = SamplingParams(
        guided_decoding=GuidedDecodingParams(
            json={"type": "object"}
        )
    )

    with pytest.raises(ValueError, match="guided_decoding_backend"):
        llm._check_arguments(
            prompt_len=1,
            sampling_params=sampling_params,
            is_gen_only=False,
        )

This test lives in:

tests/unittest/llmapi/test_sampling_params.py

Against unmodified current main, the regression test fails with:

Failed: DID NOT RAISE <class 'ValueError'>

The request contains an active guided-decoding constraint, but the missing backend is not rejected before execution.

I also traced the request path and confirmed that guided decoding is only installed when guided_decoding_backend is configured. Without a backend configured, the request can continue without the requested constraint being enforced.

I did not run full GPU inference locally because this host does not have an NVIDIA GPU. I ran the regression test inside NVIDIA's TensorRT-LLM release container built from the exact same source commit.

Container used:

nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232

Expected behavior

If a request includes an active guided-decoding constraint but the LLM/server was started without guided_decoding_backend, the request should be rejected before executor submission with a clear validation error.

Requests that do not use guided decoding should continue to work normally.

An empty/no-op GuidedDecodingParams() should also keep its current behavior and should not be rejected.

actual behavior

The request is currently accepted even when guided_decoding_backend=None.

BaseLLM._check_arguments() does not reject the mismatch, so an active guided-decoding constraint can continue toward execution even though no guided decoder was configured.

On unmodified current main, the regression test fails with:

Failed: DID NOT RAISE <class 'ValueError'>

This means the structured-output request is not being rejected at the request-vs-engine validation boundary when the required backend capability is missing.

additional notes

I tested a small fix locally in BaseLLM._check_arguments() that rejects an active guided-decoding request when no guided_decoding_backend is configured.

The check reuses the existing SamplingParams._get_guided_decoding_params() behavior instead of duplicating guide detection. That keeps the current behavior for:

  • normal requests with no guided-decoding constraint
  • empty/no-op GuidedDecodingParams()
  • requests where a guided-decoding backend is configured

Validation on the patch:

  • regression + control cases: 4 passed
  • full tests/unittest/llmapi/test_sampling_params.py: 56 passed
  • TensorRT-LLM pre-commit checks: passed

The separate OpenAI json_schema wrapper issue mentioned in NVIDIA/TensorRT#4855 is already fixed on current TensorRT-LLM main, so this issue is only about the missing-backend case.

Before submitting a new issue...
  • Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
主要言語
Python
スター
14.7k
フォーク
2.8k
平均マージ
3日 5時間
マージ済み PR(30日)
503

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/TensorRT-LLM のほかの issue

NVIDIA/TensorRT-LLM の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。