Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

MagpieTTS long-form: a 4–5 character sentence fails the whole request ("longform history context cache is too short")

オープン
#61 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 3 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
52/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
cpp

調査の方向性

Start by locating the MagpieTTS long-form synthesis path and the history-context cache check that reports “need 20 token(s)”; the issue does not name source files or tests. Reproduce with the provided command and compare long-form behavior for very short sentences versus the workaround. Done means short sentences no longer fail the request, while long-form output is not silently truncated.

索引モデルが issue の本文から書いたものです。

説明

Version: nemo-speech 0.1.0 and 0.2.0 (nemo-speech-0.{1,2}.0-linux-x86_64-cuda
release tarballs; the same input fails identically on both; the nightly was not tried), model nvidia/magpie_tts_multilingual_357m@452ef560f972
(v2602.f16.gguf), codec nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps.
RTX 3090, driver 570.153.02, Linux 6.8.

What happens: when the input is long enough for long-form mode to engage,
a single very short sentence anywhere in it makes synthesis fail:

[nemo-speech] synthesize session started
longform history context cache is too short: need 20 token(s), have 18
nemo-speech synthesize: MagpieTTS synthesis failed
[nemo-speech] synthesize session failed (exit code 1)

nemo-speech serve returns HTTP 500 "MagpieTTS synthesis failed" for the same
input. It is deterministic: the same text fails every time, with any voice.

Repro: a filler sentence repeated four times, then one short sentence,
then the filler four times again (~900 characters):

F="The service kept running normally while the operators reviewed the dashboards and the alert history in detail."
F4="$F $F $F $F"
echo "$F4 Okay. $F4" > t.txt
nemo-speech synthesize -i t.txt -o o.wav --force --seed 1
# -> exit 1, "need 20 token(s), have 18"

Which sentences fail: the same template, with only the middle sentence
changed:

middle sentence chars result
Why? 4 fails, "have 12"
Yes. 4 fails, "have 12"
Okay. 5 fails, "have 18"
No way. 7 ok
Why not? 8 ok
12 other sentences of 11–20 chars all ok

The same short sentence passes in a short input, where long-form mode does not
engage. In a real 48-chunk text (each chunk ≤ 800 chars), 47 rendered and one
failed. It contained "Why? It ran out of disk space. Why? Logs were not
rotated. Why? …".

--tts.longform options: on behaves like auto (fails). off avoids
the error but truncates silently: an 800-char chunk that gives 53 s of audio
in long-form came back as 23 s, and 31 s with --steps 3000.

Expected: a short sentence is merged with its neighbour, or padded, so that
the history context reaches the minimum. It should not fail the request.

Workaround on our side: before sending text to Magpie, the client attaches
every sentence shorter than 8 characters to the previous sentence (or to the
next one when it comes first): "down. Why?" becomes "down, why?". With that,
the 48-chunk text renders 48 of 48.

主要言語
C++
スター
150
フォーク
32
平均マージ
7日 17時間
マージ済み PR(30日)
8

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/NeMo-Speech.cpp のほかの issue

NVIDIA/NeMo-Speech.cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。