Audio pipeline: reused narration silently dropped + voice-clip seam collisions — both pass every check
Maintainers usually reply within 1 day
@miguel-heygen is already working on this.
Since Jul 27, 2026.
Assessment
This issue has not been assessed yet.
Description
Filing note (please read): This is a tool-level / systemic report, not a "my project broke" report. Everything below is written so the fix is a guardrail in the pipeline itself, so that no HyperFrames project (mine or anyone's) can ship a silent or seam-glitched narration track again. Two related defects in the audio assembly + validation path are grouped because they share one theme: the pipeline treats audio as best-effort and never verifies it. They can be split into two issues if you prefer — say the word.
Describe the bug
Two independent, reproducible defects in the audio assembly + validation pipeline. Both are silent failures — the render "succeeds", every gate is green, and the problem is only discoverable by listening or by manually probing the output with ffmpeg.
-
Defect A — Narration can be dropped from the render with zero warnings. When narration audio is supplied out-of-band (reused / pre-generated clips, i.e. the documented
audio_reusepath) the WAVs exist on disk but are never wired into the composition, because theindex.htmlassembler builds<audio>tags from an audio ledger (audio_meta.json/audio_engine_meta.json→voices[]) that the reuse path leaves empty. Result: a fully rendered video with digitally silent narration that passeshyperframes checkat 0 errors / 0 warnings. -
Defect B — Adjacent narration clips collide at scene seams. When per-scene voice clips are placed at each scene's visual start time, but scenes crossfade-overlap, consecutive voice clips either overlap (speech-on-speech) or butt-join with zero edge silence, producing audible word-jumps (overlap) and abrupt cuts (hard splice) — even though each source clip is clean and uninterrupted.
Neither defect is specific to my assets, my script, my theme, or my machine. Both are structural: A is a data-flow gap (audio produced but never registered), B is a placement-model gap (audio timeline inherits the visual crossfade timeline). And crucially, there is no audio validation gate anywhere in the pipeline, so both ship green.
Link to reproduction
I don't have a hosted public repo, but here is a deterministic minimal-repro recipe from a clean project (happy to convert to a hosted repo if that's required to triage):
# Defect A — silent narration via the reuse path
npx hyperframes init audio-repro --non-interactive --example blank
cd audio-repro
# Simulate the documented "reuse pre-generated audio" path:
# - drop 2+ narration WAVs on disk (e.g. audio/frames/frame-01.wav, frame-02.wav)
# - do NOT let the audio engine populate voices[] (this is what reuse does today)
mkdir -p audio/frames
# (copy any 2 short speech WAVs into audio/frames/)
# Build/assemble so index.html is generated from the (empty) voices[] ledger.
npx hyperframes render . -o out.mp4
# EXPECTED by user: narration audible. ACTUAL: track is digital silence.
ffmpeg -i out.mp4 -af volumedetect -f null - 2>&1 | grep mean_volume
# -> mean_volume: ~ -60 to -91 dB (silence), yet `hyperframes check` = 0 errors.
# Defect B — seam collisions when voice starts == visual scene starts
# In a multi-scene composition that uses scene crossfades, place each scene's
# narration <audio data-start> at the scene's visual start time. Because scenes
# overlap during the crossfade, clip N and clip N+1 overlap or butt-join.
The key point isn't my files — it's that the reuse path and the seam-placement model are what ship in the tool today, so any project that (A) reuses audio or (B) crossfades scenes with per-scene VO will hit this.
Steps to reproduce
Defect A (silent narration):
- Start a project and use the reuse / pre-generated audio path (
audio_reuse: true) instead of generating TTS inline. - Assemble →
index.htmlis built fromvoices[], which the reuse path never populated (voices: []). npx hyperframes render .→ succeeds.npx hyperframes check→ 0 errors, 0 warnings (passes WCAG/layout/motion/fonts).- Play the MP4 → narration is silent (only SFX, if any, are audible).
Defect B (seam collisions):
- Build a multi-scene composition where scenes crossfade (adjacent
.clips overlap by ~0.3–0.6s). - Place each scene's narration
<audio data-start="…">at the visual scene start. - Render and listen at scene boundaries → word-jumps where clips overlap, abrupt cuts where they butt-join with no breath/edge silence.
Expected behavior
- A: If narration audio is expected (script present, or WAVs on disk under the audio dir), the render must either include it or fail loudly. "Audio produced but not wired" should be impossible to ship — the assembler should reconcile
voices[]from disk, and a validation gate should assert the track is not silent. - B: Narration placement should own its own timeline independent of the visual crossfade timeline: consecutive clips must not overlap, and every clip edge should carry a tiny fade / minimum breath so splices are inaudible. Clean source clips should stay clean end-to-end.
- Both:
hyperframes checkshould cover audio (presence, loudness, expected-vs-actual voice count, seam sanity), the same way it covers contrast/layout/motion today.
Actual behavior
- A: Rendered MP4 has a valid AAC stereo track that is digital silence for narration — full-file
mean_volume ≈ -61.5 dB; every 5s probe window≈ -91.0 dB.grep "audio/frames" index.html→ 0 matches (the WAVs were never referenced).audio_meta.json/audio_engine_meta.json→"voices": [],"tts_provider": null.hyperframes check→ 0 errors / 0 warnings / 40/40 WCAG AA. Sparse SFX (the only tags that were in the ledger) produced amax_volume ≈ -14.4 dBspike that made the AAC track look "healthy" at a glance and masked the silent narration.- After manually injecting the 28 missing
<audio>tags intoindex.html, re-render → narration present atmean ≈ -21 dB, peaks -3…-6 dB. So the audio itself was always fine; it was orphaned from the composition.
- After manually injecting the 28 missing
- B: A null-test (local-align + subtract rendered vs. source) shows a 32–36 dB residual drop at every suspect point → the renderer reproduces the source faithfully (no dropped words, no corruption). The audible glitches come entirely from placement: measured 7 true speech-on-speech overlaps (effective gap negative, down to −76 ms) and ~9 hard butt-joins (≈0 gap with 0.000 s edge silence on the trimmed TTS). Overlap → "word jump"; butt-join → "abrupt cut".
Environment
hyperframes doctor
✓ Version 0.7.70 (latest)
✓ Node.js v26.0.0 (darwin arm64)
✓ CPU 8 cores · Apple M1 Pro @ 2400MHz
✓ Memory 16.0 GB total
✓ FFmpeg ffmpeg 8.0 at /opt/homebrew/bin/ffmpeg
✓ FFprobe ffprobe 8.0 at /opt/homebrew/bin/ffprobe
✓ Chrome bundled headless-shell (mac_arm 152)
✓ whisper-cpp /opt/homebrew/bin/whisper-cli
✓ TTS (Kokoro) deps installed
(Reproduced with npx [email protected]. Not machine-specific — the data-flow and placement gaps are platform-independent.)
Additional context
Root cause (why these are systemic, not project-specific)
Defect A — the audio ledger is the single source of truth, and the reuse path doesn't write to it.
The index.html assembler emits <audio data-start=…> tags from voices[]. Inline TTS populates voices[] as a side effect; the reuse / pre-generated path writes WAVs to disk but never registers them. Empty ledger → zero narration tags → silent render. This is a decoupling bug: audio production and audio registration are not atomic, so any out-of-band audio is invisible to assembly. Every project that reuses audio is exposed.
Defect B — the voice timeline inherits the visual crossfade timeline.
Per-scene VO is anchored to each scene's visual start, but scenes overlap during crossfades, so effective per-clip air-time is (scene_len − crossfade). Place full-length clips at overlapping starts and adjacent clips must collide. Hard-trimmed TTS (0 ms lead/tail silence) removes the only cushion that would hide a butt-join. Both conditions are defaults of the pipeline, so this reproduces anywhere scenes crossfade.
Cross-cutting gap — no audio validation gate. check validates layout, motion, contrast, and fonts, but never audio. A fully silent 9-minute video was "green." This is why both defects ship: nothing in the pipeline ever listens.
Proposed fixes (durable, tool-level)
A — make audio impossible to lose:
- Reconcile
voices[]from disk before assembly. Scan the audio dir (audio/frames/*.wavor equivalent); for any file not in the ledger, backfill an entry (frame,file,start_s,duration_sfromffprobe). Assemble from the reconciled ledger so reused audio is always wired. - Make produce+register atomic. Whether TTS is inline or reused, writing a WAV and appending its
voices[]entry should be one step (or a post-step reconciler runs automatically).
B — give narration its own seam-safe timeline:
3. De-overlap placement. Compute voice starts on an audio timeline, not the visual one: never let clip N+1 start before clip N ends; if a scene's air-time is shorter than its clip, push the next start (or shorten the crossfade for that boundary).
4. Micro-fades + minimum breath on every edge. Apply ~10–12 ms fade-in/out to each clip and enforce a small minimum inter-clip gap so butt-joins are inaudible while A/V sync is preserved.
5. (Optional convenience) Offer a pre-baked single narration track: de-overlap + edge-fade the clips into one file, replacing N per-scene <audio> tags with one — eliminates seam math at assembly time.
Cross-cutting — add an audio gate to check / render:
6. Loudness assertion: run ffmpeg volumedetect post-render; fail if mean_volume < -50 dB (whole file) or if any expected-narration window reads silence.
7. Count assertion: compare expected voice-clip count (from script/ledger) vs. <audio> tags actually in index.html; fail on mismatch.
8. Seam assertion: flag any two narration clips whose placement overlaps or whose effective gap < ~60 ms with ~0 edge silence.
Guardrail summary
| # | Guardrail | Catches | Where |
|---|---|---|---|
| 1 | Reconcile voices[] from disk |
Orphaned reused audio (Defect A) | assemble |
| 2 | Atomic produce+register | Ledger desync at the source (Defect A) | TTS / reuse |
| 3 | De-overlap voice placement | Speech-on-speech word-jumps (Defect B) | assemble |
| 4 | Edge micro-fades + min breath | Abrupt butt-join cuts (Defect B) | assemble |
| 5 | Loudness gate (volumedetect) |
Silent / near-silent output (A) | check / render |
| 6 | Voice-count gate | Missing narration tags (A) | check |
| 7 | Seam gate | Overlaps / hard splices (B) | check |
Acceptance criteria
- Reusing pre-generated audio results in wired, audible narration or a hard failure — never a silent success.
-
hyperframes checkfails on a silent/near-silent track, on expected-vs-actual voice-count mismatch, and on seam collisions. - With crossfaded scenes + per-scene VO, adjacent clips never overlap and every edge has a fade/min-breath; clean source stays clean.
- A green
checkguarantees audio is present and seam-safe.
Diagnostic method (reusable, for whoever triages)
- Silence probe:
ffmpeg -i out.mp4 -af volumedetect -f null -(whole file) + windowed-ss/-tscans to find silent regions. - Faithfulness null-test (rules out corruption): build an "ideal" track from raw clips via
ffmpeg adelay+amix; local-align (FFT cross-correlation on amplitude envelopes) and subtract rendered vs. ideal per suspect region — ≥12–15 dB residual drop ⇒ faithful copy (glitch is placement, not the renderer). - Seam math:
effective_gap(N) = placed_gap(N,N+1) + tail_silence(N) + lead_silence(N+1);< 0= overlap,0–~0.12 s= abrupt splice.
Minor notes
- Source narration is mono/24 kHz; output up-mixed to stereo/48 kHz cleanly (not a bug).
bgm: null— no music bed; a light bed under VO would also soften seams but isn't the fix.
- Dominant language
- TypeScript
- Stars
- 54.1k
- Forks
- 4.9k
- Avg merge
- 7h 29m
- Merged PRs (30d)
- 778
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from heygen-com/hyperframes
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
heygen-com/hyperframes#5027 ·
Maintainers usually reply within 1 day
-
fix(producer): propagate useGpu to HDR layered streaming encoderPossibly taken @Monster-GM claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 87/100
heygen-com/hyperframes#5002 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
heygen-com/hyperframes#4702 · 1 comment · 1 reaction ·
Maintainers usually reply within 1 day
-
Studio catalog prompt editor has no accessible namePossibly taken @lorenzozanee claimed this 11 days ago. Openbug difficulty/easy triage/ready
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
heygen-com/hyperframes#4384 ·
Maintainers usually reply within 1 day
-
lint: validate composition variables declared on supported root elementsMay be free again A pull request for this issue was closed without being merged. Openbug difficulty/easy triage/ready
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
heygen-com/hyperframes#4383 ·
Maintainers usually reply within 1 day
All issues in heygen-com/hyperframes
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
bug:new
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
callstackincubator/simlock#350 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
openwatersio/maritime-zones#33 ·
Maintainers usually reply within 1 day
-
Booking email verification fails for plus aliases with impersonation protection enabledPossibly taken @kankadev claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
calcom/cal.diy#30293 · 1 comment ·
Maintainers usually reply within 5 days
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
AOSSIE-Org/DebateAI#611 ·
Maintainers usually reply within 3 days