NVIDIA-NeMo/Speech

Speed up RNNT model inference using TRT

Ouverte

#14 531 ouverte le 20 août 2025

 (1 commentaire) (0 réaction) (1 personne assignée)Python (3 563 forks)github user discovery
ASRcommunity-requesthelp wantedwaiting-on-customer

Métriques du dépôt

Stars
 (18 175 étoiles)
Métriques de merge PR
 (Aucune PR mergée en 30 j)

Description

Hi,

I previously trained an RNNT model and now want to accelerate it by converting it to TensorRT. I’ve exported the model to ONNX and have encoder.onnx and decoder.onnx.

I’m using the TensorRT 25.03 Docker image and trtexec to convert the models. The decoder works fine with --fp16, but when I use --fp16 for the encoder, some outputs return NaN and the results are incorrect.

Has anyone encountered this issue or knows how to fix it?

Are there any methods to accelerate RNNT model inference?

Guide contributeur