T2VA
文字生成影片
文字提示詞
開放式創意、廣告和電影化場景。
寫清初始構圖、一個主動作、鏡頭路徑和聲音。
開啟這個工作流程01 / MODE PICKER
從這裡開始。模式選錯後,模型還沒讀到鏡頭指令,輸入之間就已經發生衝突。
T2VA
文字提示詞
開放式創意、廣告和電影化場景。
寫清初始構圖、一個主動作、鏡頭路徑和聲音。
開啟這個工作流程I2VA
首圖+提示詞
控制人物身份、產品外形和構圖。
圖片負責靜態畫面,提示詞重點說明接下來如何運動。
開啟這個工作流程FL2VA
首圖+尾圖+提示詞
狀態轉變和精確落到最終畫面。
只設計一條可達運動路徑,並保持兩幀構圖相容。
開啟這個工作流程L2VA
尾圖+提示詞
Logo 揭示、產品尾卡和預設結尾。
描述合理的前置狀態,以及元素如何逐步收束到尾圖。
開啟這個工作流程Ref2VA
圖片、影片、音訊+提示詞
遷移身份、產品、動作、運鏡或聲音。
為每個素材指定一個明確職責,並寫清必須保留什麼。
開啟這個工作流程CONTROL RULES
先滿足這些規則,再補電影化細節。輸入互相衝突時,堆更多詞不會解決問題。
輸入類型決定提示詞應該控制什麼,不要給所有工作流程套同一個萬用公式。
短影片只安排一個清楚事件,只有畫面資訊真正變化時才增加鏡頭。
拆開身份、產品、動作、鏡頭和聲音職責,避免參考素材互相爭奪控制權。
把模糊否定改成明確狀態:固定鏡頭、閉合嘴唇、完全一致的臉、服裝保持不變。
02 / PROMPT BUILDER
Builder 會隨 H3 模式切換對齊指令、參考標籤和 API content role。
生成模式
10s. A matte-red portable speaker remains geometrically consistent. A rim light traces the silhouette, then the speaker settles into a clean hero frame. Minimal dark studio, soft haze, precise reflections and generous negative space. The camera performs a Push In with small amplitude at slow speed. One soft dial click, restrained room tone and a synchronized low-frequency pulse. Minimal electronic percussion, moderate tempo, ending cleanly on the hero frame.輸出會隨欄位即時更新;Prompt 建議保留英文,對白與畫面文字使用目標語言。
462/700003 / OFFICIAL ANATOMY
把欄位當成思考框架:先寫畫面和鏡頭,再寫畫內聲音,最後寫觀眾可聽配樂。
T2VA / I2VA / FL2VA / L2VA
{I2VA / FL2VA / L2VA alignment instruction when applicable}
integrated_multimodal_description:
[Shot 1] {initial composition}. {subject action}.
[Shot 2] At 00:04.000, {new visual information}.
overall_soundscape:
{ambience + physical sounds}
non_diegetic_music:
{instrumentation + tempo + dynamics}; or N/ARef2VA
subject_definitions:
<Subject 1> is the product shown in <Picture 1>; preserve its geometry, material, color and logo.
<Video 1> is the camera-path and pacing reference.
summary:
[reference generation] {creative goal using Subject 1}
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - {exact attributes}
<Video 1> (camera and pacing structure): weak_reference - {relationship}
detailed_description:
{style sentence}
[Shot 1] <Subject 1> ...
overall_soundscape:
...
non_diegetic_music:
...tracking shot + medium amplitude + slow speed(S1) says <d>[Chinese] 原句</d>instrumentation + tempo + dynamics; or N/A04 / REFERENCE ROLES
清楚的職責映射比增加更多參考素材更有用;必須保持的細節要在文字中再寫一次。
<Subject 1>身份:臉、髮型、服裝
<Subject 2>來自 Picture 2 的產品:幾何、材質、顏色
<Picture 3>具體首幀、關鍵幀或尾幀錨點
<Video 1>動作:步行節奏與鏡頭路徑
<Audio 1>聲音:僅參考音色
保持條件
In Video 1, change only the original package to the exact product from Picture 1.
Preserve the actor identity, hands, timing, camera path, background and lighting.
Keep the product geometry, color and logo from Picture 1 unchanged.PROMPT RECIPES
這些是原創教學範本,不代表必然重現;每次只調整一個變數。
10s premium product film. A matte-red speaker in clean negative space. Slow push-in as rim light traces the exact silhouette; cut once to a macro texture shot, then settle into a stable hero frame. One soft dial click, restrained room tone, minimal electronic percussion.
使用這個 PromptPicture 1 anchors 0.00s and Picture 2 anchors 8.00s. Use one continuous shot. The hand pulls the ribbon and the paper unfolds in observable stages. Every moving element gradually settles into the exact composition, lighting and hand position of Picture 2.
使用這個 PromptIn Video 1, change only the package to the exact product from Picture 1. Preserve actor identity, hands, timing, camera path, background, lighting and all other motion. Keep product geometry, color and logo unchanged.
使用這個 Prompt05 / TROUBLESHOOTING
每次重試只改變一個控制變數,才能判斷哪條指令真正改善了結果。
| 症狀 | 可能原因 | 下一步修改 |
|---|---|---|
| 人物或產品漂移 | 參考職責不明確,或同時承擔風格和動作。 | 用一張乾淨圖片只負責身份,並列出必須保持的具體特徵。 |
| 鏡頭運動衝突 | 影片參考和文字描述了不同鏡頭路徑。 | 讓參考影片負責路徑,文字只補充速度或幅度。 |
| 對白過快或聲畫錯位 | 台詞過長、speaker 變化,或沒有綁定時間節點。 | 縮短台詞、保持同一 speaker ID,並放進對應鏡頭。 |
| 首尾幀突然跳變 | 兩幀的機位、尺度差異過大,或獨立運動物件太多。 | 從首幀編輯出尾幀,只改變一個可達狀態,並描述一條連續路徑。 |
06 / CONTRACT BOUNDARIES
所有時長、解析度、參數和 model ID 都必須標註具體由哪個平台提供。
官方 API 時長
4–15s
官方解析度
768P / 2K
首尾幀比例
adaptive
文字生成影片比例
21:9 → 9:16
請求 schema、model string 與能力邊界以官方文件為準。
官方 API 指南本站可用模式、控制項、價格與限制以目前工作台顯示為準。
開啟工作台按 T2V、首尾幀或參考生成選擇對應 workflow 與權重。
ComfyUI 指南FAQ
模式、參考素材、對白和平台合約都應該在風格潤飾前確定。
純文字創意用 T2VA;首幀決定構圖用 I2VA;首尾都重要用 FL2VA;預設結尾用 L2VA;需要圖片、影片或音訊控制特定屬性時用 Ref2VA。
不要假設所有 Provider 都有該欄位。改寫成正向保持條件,例如「鏡頭保持固定」「保留完全一致的臉和服裝」「嘴唇保持閉合」。
保持 speaker ID 一致,標註對白語言,讓台詞長度匹配時長,並把對白、環境聲、物理音效和觀眾可聽配樂分開。
不一樣。權重、輸入模式、時長、解析度和參數都可能不同,必須以目前使用平台的合約為準。
來源與核驗
內容最後核驗:2026 年 8 月 11 日。頁面範本是依據所連結結構原創編寫的教學材料。