Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug]: guided response_format is silently ignored without guided_decoding_backend

Open Beginner friendly
#19,635 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
api, backend

Research direction

Start with BaseLLM._check_arguments() and read the related helpers in tests/unittest/llmapi/test_sampling_params.py. Run the named regression test first, then use the reported control cases to verify that active guided decoding without guided_decoding_backend is rejected while no-op and configured-backend requests retain their behavior.

Written by the indexing model from the issue text.

Description

bug Inference runtime
System Info

TensorRT-LLM commit:
c76f4a856447820aa1f994b053da7580d69b657d

Container:
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232

Container environment:

  • NVIDIA Release 26.08
  • PyTorch 2.14.0a0+4fdf77b
  • Python 3.12

Host:

  • Linux x86_64
  • Docker 29.1.3
  • CPU-only host for the regression test

The container reports the same TRT_LLM_GIT_COMMIT as the source revision above.

Who can help?

No response

Information
  • The official example scripts
  • My own modified scripts
Tasks
  • An officially supported task in the examples folder (such as GLUE/SQuAD, ...)
  • My own task or dataset (give details below)
Reproduction

I ran into the remaining missing-backend case from NVIDIA/TensorRT#4855 and reproduced it against current TensorRT-LLM main:

c76f4a856447820aa1f994b053da7580d69b657d

I reduced the issue to the request-vs-engine validation path in BaseLLM._check_arguments().

The regression test I used is:

def test_check_arguments_rejects_guided_decoding_without_backend() -> None:
    llm = _TestLLM(guided_decoding_backend=None)
    sampling_params = SamplingParams(
        guided_decoding=GuidedDecodingParams(
            json={"type": "object"}
        )
    )

    with pytest.raises(ValueError, match="guided_decoding_backend"):
        llm._check_arguments(
            prompt_len=1,
            sampling_params=sampling_params,
            is_gen_only=False,
        )

This test lives in:

tests/unittest/llmapi/test_sampling_params.py

Against unmodified current main, the regression test fails with:

Failed: DID NOT RAISE <class 'ValueError'>

The request contains an active guided-decoding constraint, but the missing backend is not rejected before execution.

I also traced the request path and confirmed that guided decoding is only installed when guided_decoding_backend is configured. Without a backend configured, the request can continue without the requested constraint being enforced.

I did not run full GPU inference locally because this host does not have an NVIDIA GPU. I ran the regression test inside NVIDIA's TensorRT-LLM release container built from the exact same source commit.

Container used:

nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc29.dev202609250232

Expected behavior

If a request includes an active guided-decoding constraint but the LLM/server was started without guided_decoding_backend, the request should be rejected before executor submission with a clear validation error.

Requests that do not use guided decoding should continue to work normally.

An empty/no-op GuidedDecodingParams() should also keep its current behavior and should not be rejected.

actual behavior

The request is currently accepted even when guided_decoding_backend=None.

BaseLLM._check_arguments() does not reject the mismatch, so an active guided-decoding constraint can continue toward execution even though no guided decoder was configured.

On unmodified current main, the regression test fails with:

Failed: DID NOT RAISE <class 'ValueError'>

This means the structured-output request is not being rejected at the request-vs-engine validation boundary when the required backend capability is missing.

additional notes

I tested a small fix locally in BaseLLM._check_arguments() that rejects an active guided-decoding request when no guided_decoding_backend is configured.

The check reuses the existing SamplingParams._get_guided_decoding_params() behavior instead of duplicating guide detection. That keeps the current behavior for:

  • normal requests with no guided-decoding constraint
  • empty/no-op GuidedDecodingParams()
  • requests where a guided-decoding backend is configured

Validation on the patch:

  • regression + control cases: 4 passed
  • full tests/unittest/llmapi/test_sampling_params.py: 56 passed
  • TensorRT-LLM pre-commit checks: passed

The separate OpenAI json_schema wrapper issue mentioned in NVIDIA/TensorRT#4855 is already fixed on current TensorRT-LLM main, so this issue is only about the missing-backend case.

Before submitting a new issue...
  • Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
3d 5h
Merged PRs (30d)
504

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/TensorRT-LLM

All issues in NVIDIA/TensorRT-LLM

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.