[Bug]: guided response_format is silently ignored without guided_decoding_backend
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 82/100
Research direction
Start with BaseLLM._check_arguments() and read the related helpers in tests/unittest/llmapi/test_sampling_params.py. Run the named regression test first, then use the reported control cases to verify that active guided decoding without guided_decoding_backend is rejected while no-op and configured-backend requests retain their behavior.
Written by the indexing model from the issue text.
Description
System Info
TensorRT-LLM commit:
c76f4a856447820aa1f994b053da7580d69b657d
Container:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232
Container environment:
- NVIDIA Release 26.08
- PyTorch 2.14.0a0+4fdf77b
- Python 3.12
Host:
- Linux x86_64
- Docker 29.1.3
- CPU-only host for the regression test
The container reports the same TRT_LLM_GIT_COMMIT as the source revision above.
Who can help?
No response
Information
- The official example scripts
- My own modified scripts
Tasks
- An officially supported task in the
examplesfolder (such as GLUE/SQuAD, ...) - My own task or dataset (give details below)
Reproduction
I ran into the remaining missing-backend case from NVIDIA/TensorRT#4855 and reproduced it against current TensorRT-LLM main:
c76f4a856447820aa1f994b053da7580d69b657d
I reduced the issue to the request-vs-engine validation path in BaseLLM._check_arguments().
The regression test I used is:
def test_check_arguments_rejects_guided_decoding_without_backend() -> None:
llm = _TestLLM(guided_decoding_backend=None)
sampling_params = SamplingParams(
guided_decoding=GuidedDecodingParams(
json={"type": "object"}
)
)
with pytest.raises(ValueError, match="guided_decoding_backend"):
llm._check_arguments(
prompt_len=1,
sampling_params=sampling_params,
is_gen_only=False,
)
This test lives in:
tests/unittest/llmapi/test_sampling_params.py
Against unmodified current main, the regression test fails with:
Failed: DID NOT RAISE <class 'ValueError'>
The request contains an active guided-decoding constraint, but the missing backend is not rejected before execution.
I also traced the request path and confirmed that guided decoding is only installed when guided_decoding_backend is configured. Without a backend configured, the request can continue without the requested constraint being enforced.
I did not run full GPU inference locally because this host does not have an NVIDIA GPU. I ran the regression test inside NVIDIA's TensorRT-LLM release container built from the exact same source commit.
Container used:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232
Expected behavior
If a request includes an active guided-decoding constraint but the LLM/server was started without guided_decoding_backend, the request should be rejected before executor submission with a clear validation error.
Requests that do not use guided decoding should continue to work normally.
An empty/no-op GuidedDecodingParams() should also keep its current behavior and should not be rejected.
actual behavior
The request is currently accepted even when guided_decoding_backend=None.
BaseLLM._check_arguments() does not reject the mismatch, so an active guided-decoding constraint can continue toward execution even though no guided decoder was configured.
On unmodified current main, the regression test fails with:
Failed: DID NOT RAISE <class 'ValueError'>
This means the structured-output request is not being rejected at the request-vs-engine validation boundary when the required backend capability is missing.
additional notes
I tested a small fix locally in BaseLLM._check_arguments() that rejects an active guided-decoding request when no guided_decoding_backend is configured.
The check reuses the existing SamplingParams._get_guided_decoding_params() behavior instead of duplicating guide detection. That keeps the current behavior for:
- normal requests with no guided-decoding constraint
- empty/no-op
GuidedDecodingParams() - requests where a guided-decoding backend is configured
Validation on the patch:
- regression + control cases: 4 passed
- full
tests/unittest/llmapi/test_sampling_params.py: 56 passed - TensorRT-LLM pre-commit checks: passed
The separate OpenAI json_schema wrapper issue mentioned in NVIDIA/TensorRT#4855 is already fixed on current TensorRT-LLM main, so this issue is only about the missing-backend case.
Before submitting a new issue...
- Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
- Dominant language
- Python
- Stars
- 14.7k
- Forks
- 2.8k
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 504
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/TensorRT-LLM
-
Customized kernels
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
NVIDIA/TensorRT-LLM#19564 · 1 comment ·
Maintainers usually reply within 1 day
-
AutoDeploy bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TensorRT-LLM#19491 · 2 comments ·
Maintainers usually reply within 1 day
-
Customized kernels
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#19240 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 62/100
NVIDIA/TensorRT-LLM#19225 ·
Maintainers usually reply within 1 day
-
Infra
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#19067 ·
Maintainers usually reply within 1 day
All issues in NVIDIA/TensorRT-LLM
Similar issues
-
namespace operations
Difficulty 1/5 Under an hour Newbie friendliness 82/100
EclipseFdn/open-vsx.org#13573 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
collective/icalendar#1854 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
rancher/rancher-ai-agent#412 ·
Maintainers usually reply within 6 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
TUDelftGeodesy/DePSI#134 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
HenriquesLab/rxiv-maker#335 ·