Add source-audio-synchronized performance shots and provider-aware music-video timing
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
この issue はまだ評価されていません。
説明
Problem
Music Video can cut generated clips over a master song, but cannot currently instruct the Grok backend to synchronize a visible singer to that recording. Prompting alone cannot close the missing audio-conditioning contract.
Verified evidence (2026-09-27)
server/services/videoGen/grok.js:buildGrokVideoPromptsends source image, motion prose and duration through Grok Build CLI image_to_video.generateVideoaccepts no audio input or timed lyric cues.server/lib/grokVideoClip.jslimits that CLI lane to 6/10 seconds based on historical rendered measurements. Do not equate this CLI contract with the newer REST API.server/services/videoGen/submitJob.jshosted dispatch passes prompt/image/duration but not song audio.client/src/hooks/useMusicVideoSceneMedia.jsuses a saved grokDuration rather than deriving the generated coverage from each scene. Its separate local audio-reactive guard explicitly prohibits singing/lip sync.server/services/musicVideo/render.jsconcatenates video over the master track and may loop a short source; this aligns cuts, not mouth movements.- xAI reference-to-video docs describe timed keyframes (1/3-second grid) and preset voice identities. Own audio-file voice references require trusted-partner access. These docs do not establish exact song waveform/phoneme preservation: https://docs.x.ai/developers/model-capabilities/video/reference-to-video . Do not advertise existing Grok CLI as supporting these REST features.
- A documented source-audio route exists at https://fal.ai/models/minimax/h3-max/lip-sync/image-to-video/api : image_url + audio_url, optional transcription, at least 5 seconds of audio, truncation beyond 14.8 seconds, output duration follows clipped audio. Actual singing quality is unvalidated.
Chosen implementation
Add an explicit performance/lip-sync scene mode distinct from cutaway and environmental audio-reactive scenes. Initially integrate the documented fal H3 lip-sync route via capability-specific payloads, reusing the fal queue/download lifecycle once #8968 lands. This is not fulfilled by exposing the current Hailuo image-to-video model or merely changing its model ID.
Create an immutable shot instruction record from project audio revision/hash, absolute song interval, clip-relative cue times, reference image/selected character, intended performance, model capability snapshot, target edit duration and generated coverage. Slice the exact master audio segment with ffmpeg at the public generation boundary. Optional vocal stems must preserve the original timebase; retain the original master in final assembly. Use transcription guidance only where supported.
For source-audio generation, plan within provider limits; prevent silent 14.8-second truncation. Short shots may generate a supported-length contextual audio window with an explicit edit in-point, then trim video/audio at identical offsets. Split longer passages on musical/lyric boundaries. Never loop or time-stretch a singing take to fill a slot. Cancel/failure/restart must not duplicate paid submissions or replace selected takes.
For Grok CLI cutaways, derive covering 6/10-second requests from shot duration (reuse nearestGrokDuration), split spans over 10 seconds, include relative motion cues in prose, and validate actual output duration before trimming. Describe internal gesture timing as approximate. Reject performance mode for Grok until that actual adapter has a verified exact-source-audio capability; never silently generate an unrelated voice/audio track. Preserve the environmental-motion no-performance guard in its current mode.
Affected: musicVideo schemas/stores/planner/render, scene media hook and settings/SceneCard UI, videoGen prepareParams/hostedSubmission/submitJob/fal, grokVideoClip, media-job persistence. Coordinate with #8964 lyric planning, #8965 take selection and #8966 composition through additive contracts; do not edit those agents' branches or broaden their assigned issues.
Blocked by #8968
Acceptance criteria
- Rendered UI and route tests show that performance mode requires a verified source-audio provider; Grok's existing lane is labelled as cutaway-only for this purpose.
- A synthetic song excerpt at a nonzero start has correct sample-window content, cue offsets and frame coverage through generation payload, trim and final timeline. Test min/max duration boundaries and cancellation/restart without paid resubmission.
- Existing project compatibility, PostgreSQL/file-test parity and sync version gates remain intact. Credentials and execution settings remain machine-local.
- No provider calls occur at boot. Explicit generation shows provider/model and estimated or unknown cost. Upload only the user-authorized selected media; use outbound provider upload/data URIs, never expose PortOS via tunnels or public callbacks.
- With an operator-authorized test budget, inspect a real vocal sample with plosives, pauses and sustained notes; review beginning/middle/end and continuous playback against the original master. Record measured offset/drift where a valid tool exists and timecoded visual findings. Mock tests or matching durations cannot certify lip-sync quality.
- Preserve the current CLI duration limits until new actual tool capabilities are validated. REST documentation alone is not evidence that the installed CLI gained those controls.
Dispatch: model:heavy for shared audio/video timebase, provider capability and resumable paid-job semantics; effort:high for end-to-end integration and audiovisual validation.
- 主要言語
- JavaScript
- スター
- 38
- フォーク
- 32
- 平均マージ
- 24分
- マージ済み PR(30日)
- 993
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
atomantic/PortOS のほかの issue
-
area:media area:ui effort:high model:heavy plan planner:gpt-6-1-sol
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 1 日以内に返信
-
area:media area:ui blocked effort:xhigh model:heavy plan planner:gpt-6-1-sol
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
メンテナーはふだん 1 日以内に返信
-
area:media area:ui blocked effort:high model:heavy plan planner:gpt-6-1-sol
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
メンテナーはふだん 1 日以内に返信
-
area:media area:ui effort:high model:heavy plan planner:gpt-6-1-sol
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 1 日以内に返信
-
area:media area:ui blocked effort:xhigh model:heavy plan planner:gpt-6-1-sol
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 1 日以内に返信
atomantic/PortOS の issue をすべて見る
似ている issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
openlibhums/janeway#5604 ·
メンテナーはふだん 1 日以内に返信
-
[BUG] Generic OSC does not initialize OSC client on startup when "Listen for Feedback" is disabledオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
-
area/statement-execution TS conversion
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
scylladb/nodejs-rs-driver#584 ·
メンテナーはふだん 2 日以内に返信
-
新讀者走讀回報,照著一篇文章實際操作オープンdocumentation good first issue help wanted
難易度 1/5 1〜3時間 初心者へのやさしさ 92/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 69/100
メンテナーはふだん 3 日以内に返信