Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Long-form instrumental collapse into radio/station-surfing montage + hallucinated lyrics (~10–20s)

未关闭
#7 1 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
需要澄清
活跃度
活跃
领域
ai

调研方向

首先,在本地 ComfyUI 中使用 EmptyMiniMaxMusic3LatentAudio、MiniMaxMusic3TextEncode、KSampler 和 VAEDecodeAudioTiled,并使用提供的提示词以及种子 101 和 202,复现该问题。如果有可用的首选参考配置,请将 INT8 路径与其进行比较;在确定类似收音机的崩溃和幻觉般的人声是否能够复现,并记录任何已确认的缓解措施后,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Summary

When generating a 2-minute fully instrumental cyberpunk cue (strict no-vocals / no-drums / no-bass), MiniMax Music 3 starts coherently for roughly the first 10–20 seconds, then collapses into what sounds like a broken radio / station-surfing montage: abrupt genre/timbre cuts as if switching broadcast stations, compressed “over the air” texture, and hallucinated vocals/lyrics despite an empty instrumental lyrics track and an explicit no-vocals caption.

This reproduced on two independent seeds with the same prompt.

Environment

  • Model: MiniMax Music 3 (ComfyUI INT8 path: minimax_music3_dit_int8_convrot + pruned INT8 text encoder + DAV)
  • Backend: local ComfyUI (EmptyMiniMaxMusic3LatentAudio + MiniMaxMusic3TextEncode + KSampler + VAEDecodeAudioTiled)
  • GPU: NVIDIA GeForce RTX 4070 Ti (12 GB)
  • Target duration: 120 s
  • Seeds that failed the same way: 101, 202
  • Output duration observed: ~120 s each

Prompt (exact)

Lyrics

[instrumental]

Caption

Global Metadata
Genre: instrumental synthwave / cyberpunk film score. BPM: 92. Key: D minor. Neon-noir, cinematic, epic undertones — like a Blade Runner rooftop climax or a rain-soaked chase resolve. Slow-burn grandeur without turning into a pop song.
Application: a two-minute cyberpunk movie cue: melancholic, heroic, and widescreen.

Vocal Details
STRICTLY NO VOCALS. No singing, no speech, no choir, no hummed melody, no vocoder voice. Fully instrumental.

Arrangement
ONLY these three sound sources, throughout the entire piece:
1) Electric guitar — clean-to-slightly-driven leads and atmospheric chords, cinematic and expressive.
2) Saxophone — smoky, lyrical cyberpunk noir lines; occasional heroic long tones.
3) Gamelan — metallophone / gong / bronze-bar textures and interlocking patterns, used as the harmonic and rhythmic sparkle (not as drums).

ABSOLUTELY NO DRUMS of any kind (no kick, snare, hi-hat, percussion kit, electronic beats, trap hats, industrial hits).
ABSOLUTELY NO BASS of any kind (no bass guitar, no synth bass, no sub bass, no 808).
No pads-as-orchestra beds that replace the trio; keep the mix as guitar + saxophone + gamelan only. Let guitar and sax carry melody; let gamelan provide shimmering rhythmic pulse and epic metallic resonance. Build intensity through harmony, register, and density — never through drums or bass.

Observed failure mode

  1. 0–~15s: mostly on-brief — instrumental cyberpunk / noir cinematic texture consistent with guitar/sax/gamelan intent.
  2. After ~10–20s: structure breaks. The track starts behaving like channel surfing: sudden jumps between unrelated musical fragments, as if flipping between radio stations.
  3. Vocals appear anyway (spoken/sung fragments / lyric-like content) despite [instrumental] + strong no-vocals constraints.
  4. Arrangement contract (only guitar + sax + gamelan; no drums/bass) is largely abandoned after the collapse.

This does not sound like ordinary mild long-form drift (instrument fading / emotion softening). It sounds specifically like a broadcast montage / radio continuity prior that the global model falls into once the initial plan loses grip.

Notes / hypothesis

Official docs already say section tags and descriptions are not strict symbolic guarantees (tempo/key/instrumentation/lyrics/structure may mismatch). Separately, the Music 3.0 blog discusses long-form brief-drift as a problem Music 3 aims to mitigate.

I could not find documentation of this specific radio/station-surfing + lyric hallucination failure mode. Happy to provide short audio excerpts from the two failing takes if useful.

Ask

  • Is this a known failure mode (e.g. radio/broadcast material in training / long-context mode collapse)?
  • Any recommended mitigations for multi-minute strict instrumental generations (caption structure, section tags, CFG, shorter chunking, etc.)?
  • If this is unexpected on the INT8 Comfy path, is there a preferred reference config to re-test against?

Thanks!

主要语言
没有语言数据
星标
896
派生
84
PR 合并指标
30 天内没有已合并 PR

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

MiniMax-AI/MiniMax-Music3 的其他 Issue

查看 MiniMax-AI/MiniMax-Music3 的全部 Issue

相似的 Issue

更多 AI Infra & Agents Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。