vllm-project/vllm-omni

[New Model] Add TADA (Hume AI) TTS model support

Open

#2 001 ouverte le 19 mars 2026

Voir sur GitHub
 (8 commentaires) (0 réactions) (1 assigné)Python (1 067 forks)github user discovery
good first issuehelp wantednew model

Métriques du dépôt

Stars
 (4 990 stars)
Métriques de merge PR
 (Métriques PR en attente)

Description

Model

TADA by Hume AI — a speech-language model that unifies speech synthesis with text generation using 1:1 text-speech token alignment built on Llama 3.2.

Architecture

  • Encoder (HumeAI/tada-codec): encodes reference audio into aligned token sequences
  • LLM (TadaForCausalLM): autoregressive generation, available in 1B and 3B-ML (multilingual) variants
  • Generates the complete speech segment per text token in a single AR step (dynamic duration/prosody)

This would likely be a 2-stage pipeline in vllm-omni (AR generation → codec decode), similar to Qwen3 TTS / Fish Speech.

References

cc @ashok-arora — would you be interested in taking a look at this?

Guide contributeur