sgl-project/sglang

[Feature] Reduce constrained-decoding overhead in TP

Aberta

#13.809 aberto em 23 de nov. de 2025

 (2 comentários) (0 reação) (1 responsável)Python (6.216 forks)auto 404
good first issue

Métricas do repositório

Stars
 (28.442 estrelas)
Métricas de merge de PR
 (Mesclagem média 1d 13h) (1.000 fundiu PRs em 30d)

Description

Checklist

Motivation

Currently, when tensor-parallelism (TP) and constrained-decoding are enabled, each TP worker will compile the same grammar across different TP ranks. This will incur non-trivial CPU overhead (for TP $n$, the overhead is $n \times$. In fact, we only need to compile the grammar and sample the result on the first rank.

A possible implementation can be:

  1. Only apply grammar token mask on the rank 0.
  2. Broadcast the next tokens id from rank 0 when there're grammars in batch.
  3. Keep the old code path when there's no grammar in batch (i.e. no extra all-reduce).

https://github.com/sgl-project/sglang/blob/5c2915494c83f076a11afb2c3382eeb8a41f1974/python/sglang/srt/layers/sampler.py#L205-L217

Related resources

No response

Guia do colaborador