[Bug]: guided response_format is silently ignored without guided_decoding_backend
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 82/100
調査の方向性
BaseLLM._check_arguments() から始め、tests/unittest/llmapi/test_sampling_params.py にある関連ヘルパーを読みます。まず指定された回帰テストを実行し、次に報告されたコントロールケースを使って、guided_decoding_backend のないアクティブな guided decoding が拒否される一方で、no-op のリクエストと設定済みバックエンドを使用するリクエストでは動作が維持されることを確認します。
索引モデルが issue の本文から書いたものです。
説明
System Info
TensorRT-LLM commit:
c76f4a856447820aa1f994b053da7580d69b657d
Container:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232
Container environment:
- NVIDIA Release 26.08
- PyTorch 2.14.0a0+4fdf77b
- Python 3.12
Host:
- Linux x86_64
- Docker 29.1.3
- CPU-only host for the regression test
The container reports the same TRT_LLM_GIT_COMMIT as the source revision above.
Who can help?
No response
Information
- The official example scripts
- My own modified scripts
Tasks
- An officially supported task in the
examplesfolder (such as GLUE/SQuAD, ...) - My own task or dataset (give details below)
Reproduction
I ran into the remaining missing-backend case from NVIDIA/TensorRT#4855 and reproduced it against current TensorRT-LLM main:
c76f4a856447820aa1f994b053da7580d69b657d
I reduced the issue to the request-vs-engine validation path in BaseLLM._check_arguments().
The regression test I used is:
def test_check_arguments_rejects_guided_decoding_without_backend() -> None:
llm = _TestLLM(guided_decoding_backend=None)
sampling_params = SamplingParams(
guided_decoding=GuidedDecodingParams(
json={"type": "object"}
)
)
with pytest.raises(ValueError, match="guided_decoding_backend"):
llm._check_arguments(
prompt_len=1,
sampling_params=sampling_params,
is_gen_only=False,
)
This test lives in:
tests/unittest/llmapi/test_sampling_params.py
Against unmodified current main, the regression test fails with:
Failed: DID NOT RAISE <class 'ValueError'>
The request contains an active guided-decoding constraint, but the missing backend is not rejected before execution.
I also traced the request path and confirmed that guided decoding is only installed when guided_decoding_backend is configured. Without a backend configured, the request can continue without the requested constraint being enforced.
I did not run full GPU inference locally because this host does not have an NVIDIA GPU. I ran the regression test inside NVIDIA's TensorRT-LLM release container built from the exact same source commit.
Container used:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232
Expected behavior
If a request includes an active guided-decoding constraint but the LLM/server was started without guided_decoding_backend, the request should be rejected before executor submission with a clear validation error.
Requests that do not use guided decoding should continue to work normally.
An empty/no-op GuidedDecodingParams() should also keep its current behavior and should not be rejected.
actual behavior
The request is currently accepted even when guided_decoding_backend=None.
BaseLLM._check_arguments() does not reject the mismatch, so an active guided-decoding constraint can continue toward execution even though no guided decoder was configured.
On unmodified current main, the regression test fails with:
Failed: DID NOT RAISE <class 'ValueError'>
This means the structured-output request is not being rejected at the request-vs-engine validation boundary when the required backend capability is missing.
additional notes
I tested a small fix locally in BaseLLM._check_arguments() that rejects an active guided-decoding request when no guided_decoding_backend is configured.
The check reuses the existing SamplingParams._get_guided_decoding_params() behavior instead of duplicating guide detection. That keeps the current behavior for:
- normal requests with no guided-decoding constraint
- empty/no-op
GuidedDecodingParams() - requests where a guided-decoding backend is configured
Validation on the patch:
- regression + control cases: 4 passed
- full
tests/unittest/llmapi/test_sampling_params.py: 56 passed - TensorRT-LLM pre-commit checks: passed
The separate OpenAI json_schema wrapper issue mentioned in NVIDIA/TensorRT#4855 is already fixed on current TensorRT-LLM main, so this issue is only about the missing-backend case.
Before submitting a new issue...
- Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
- 主要言語
- Python
- スター
- 14.7k
- フォーク
- 2.8k
- 平均マージ
- 3日 5時間
- マージ済み PR(30日)
- 503
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TensorRT-LLM のほかの issue
-
Customized kernels
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
NVIDIA/TensorRT-LLM#19564 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
AutoDeploy bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TensorRT-LLM#19491 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
Customized kernels
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TensorRT-LLM#19240 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 62/100
NVIDIA/TensorRT-LLM#19225 ·
メンテナーはふだん 1 日以内に返信
-
Infra
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TensorRT-LLM#19067 ·
メンテナーはふだん 1 日以内に返信
NVIDIA/TensorRT-LLM の issue をすべて見る
似ている issue
-
bug server
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
sportsdataverse/sportsdataverse-py#641 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
googleapis/google-cloud-python#18532 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信