Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Add source-audio-synchronized performance shots and provider-aware music-video timing

クローズ
#8,977 コメント 4 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

この issue はまだ評価されていません。

説明

effort:high enhancement model:heavy needs-input plan

Problem

Music Video can cut generated clips over a master song, but cannot currently instruct the Grok backend to synchronize a visible singer to that recording. Prompting alone cannot close the missing audio-conditioning contract.

Verified evidence (2026-09-27)

  • server/services/videoGen/grok.js:buildGrokVideoPrompt sends source image, motion prose and duration through Grok Build CLI image_to_video. generateVideo accepts no audio input or timed lyric cues.
  • server/lib/grokVideoClip.js limits that CLI lane to 6/10 seconds based on historical rendered measurements. Do not equate this CLI contract with the newer REST API.
  • server/services/videoGen/submitJob.js hosted dispatch passes prompt/image/duration but not song audio.
  • client/src/hooks/useMusicVideoSceneMedia.js uses a saved grokDuration rather than deriving the generated coverage from each scene. Its separate local audio-reactive guard explicitly prohibits singing/lip sync.
  • server/services/musicVideo/render.js concatenates video over the master track and may loop a short source; this aligns cuts, not mouth movements.
  • xAI reference-to-video docs describe timed keyframes (1/3-second grid) and preset voice identities. Own audio-file voice references require trusted-partner access. These docs do not establish exact song waveform/phoneme preservation: https://docs.x.ai/developers/model-capabilities/video/reference-to-video . Do not advertise existing Grok CLI as supporting these REST features.
  • A documented source-audio route exists at https://fal.ai/models/minimax/h3-max/lip-sync/image-to-video/api : image_url + audio_url, optional transcription, at least 5 seconds of audio, truncation beyond 14.8 seconds, output duration follows clipped audio. Actual singing quality is unvalidated.

Chosen implementation

Add an explicit performance/lip-sync scene mode distinct from cutaway and environmental audio-reactive scenes. Initially integrate the documented fal H3 lip-sync route via capability-specific payloads, reusing the fal queue/download lifecycle once #8968 lands. This is not fulfilled by exposing the current Hailuo image-to-video model or merely changing its model ID.

Create an immutable shot instruction record from project audio revision/hash, absolute song interval, clip-relative cue times, reference image/selected character, intended performance, model capability snapshot, target edit duration and generated coverage. Slice the exact master audio segment with ffmpeg at the public generation boundary. Optional vocal stems must preserve the original timebase; retain the original master in final assembly. Use transcription guidance only where supported.

For source-audio generation, plan within provider limits; prevent silent 14.8-second truncation. Short shots may generate a supported-length contextual audio window with an explicit edit in-point, then trim video/audio at identical offsets. Split longer passages on musical/lyric boundaries. Never loop or time-stretch a singing take to fill a slot. Cancel/failure/restart must not duplicate paid submissions or replace selected takes.

For Grok CLI cutaways, derive covering 6/10-second requests from shot duration (reuse nearestGrokDuration), split spans over 10 seconds, include relative motion cues in prose, and validate actual output duration before trimming. Describe internal gesture timing as approximate. Reject performance mode for Grok until that actual adapter has a verified exact-source-audio capability; never silently generate an unrelated voice/audio track. Preserve the environmental-motion no-performance guard in its current mode.

Affected: musicVideo schemas/stores/planner/render, scene media hook and settings/SceneCard UI, videoGen prepareParams/hostedSubmission/submitJob/fal, grokVideoClip, media-job persistence. Coordinate with #8964 lyric planning, #8965 take selection and #8966 composition through additive contracts; do not edit those agents' branches or broaden their assigned issues.

Blocked by #8968

Acceptance criteria

  • Rendered UI and route tests show that performance mode requires a verified source-audio provider; Grok's existing lane is labelled as cutaway-only for this purpose.
  • A synthetic song excerpt at a nonzero start has correct sample-window content, cue offsets and frame coverage through generation payload, trim and final timeline. Test min/max duration boundaries and cancellation/restart without paid resubmission.
  • Existing project compatibility, PostgreSQL/file-test parity and sync version gates remain intact. Credentials and execution settings remain machine-local.
  • No provider calls occur at boot. Explicit generation shows provider/model and estimated or unknown cost. Upload only the user-authorized selected media; use outbound provider upload/data URIs, never expose PortOS via tunnels or public callbacks.
  • With an operator-authorized test budget, inspect a real vocal sample with plosives, pauses and sustained notes; review beginning/middle/end and continuous playback against the original master. Record measured offset/drift where a valid tool exists and timecoded visual findings. Mock tests or matching durations cannot certify lip-sync quality.
  • Preserve the current CLI duration limits until new actual tool capabilities are validated. REST documentation alone is not evidence that the installed CLI gained those controls.

Dispatch: model:heavy for shared audio/video timebase, provider capability and resumable paid-job semantics; effort:high for end-to-end integration and audiovisual validation.

主要言語
JavaScript
スター
38
フォーク
32
平均マージ
24分
マージ済み PR(30日)
993

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

atomantic/PortOS のほかの issue

atomantic/PortOS の issue をすべて見る

似ている issue

JavaScript の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。