T2VA
文生视频
文字提示词
开放式创意、广告和电影化场景。
写清初始构图、一个主动作、镜头路径和声音。
打开这个工作流01 / MODE PICKER
从这里开始。模式选错后,模型还没读到镜头指令,输入之间就已经发生冲突。
T2VA
文字提示词
开放式创意、广告和电影化场景。
写清初始构图、一个主动作、镜头路径和声音。
打开这个工作流I2VA
首图+提示词
控制人物身份、产品外形和构图。
图片负责静态画面,提示词重点说明接下来如何运动。
打开这个工作流FL2VA
首图+尾图+提示词
状态转变和精确落到最终画面。
只设计一条可达运动路径,并保持两帧构图兼容。
打开这个工作流L2VA
尾图+提示词
Logo 揭示、产品尾卡和预设结尾。
描述合理的前置状态,以及元素如何逐步收束到尾图。
打开这个工作流Ref2VA
图片、视频、音频+提示词
迁移身份、产品、动作、运镜或声音。
为每个素材指定一个明确职责,并写清必须保留什么。
打开这个工作流CONTROL RULES
先满足这些规则,再补电影化细节。输入互相冲突时,堆更多词不会解决问题。
输入类型决定提示词应该控制什么,不要给所有工作流套同一个万能公式。
短视频只安排一个清楚事件,只有画面信息真正变化时才增加镜头。
拆开身份、产品、动作、镜头和声音职责,避免参考素材互相争夺控制权。
把模糊否定改成明确状态:固定镜头、闭合嘴唇、完全一致的脸、服装保持不变。
02 / PROMPT BUILDER
Builder 会随 H3 模式切换对齐指令、参考标签和 API content role。
生成模式
10s. A matte-red portable speaker remains geometrically consistent. A rim light traces the silhouette, then the speaker settles into a clean hero frame. Minimal dark studio, soft haze, precise reflections and generous negative space. The camera performs a Push In with small amplitude at slow speed. One soft dial click, restrained room tone and a synchronized low-frequency pulse. Minimal electronic percussion, moderate tempo, ending cleanly on the hero frame.输出会随字段实时更新;Prompt 建议保留英文,对白与画面文字使用目标语言。
462/700003 / OFFICIAL ANATOMY
把字段当成思考框架:先写画面和镜头,再写画内声音,最后写观众可听配乐。
T2VA / I2VA / FL2VA / L2VA
{I2VA / FL2VA / L2VA alignment instruction when applicable}
integrated_multimodal_description:
[Shot 1] {initial composition}. {subject action}.
[Shot 2] At 00:04.000, {new visual information}.
overall_soundscape:
{ambience + physical sounds}
non_diegetic_music:
{instrumentation + tempo + dynamics}; or N/ARef2VA
subject_definitions:
<Subject 1> is the product shown in <Picture 1>; preserve its geometry, material, color and logo.
<Video 1> is the camera-path and pacing reference.
summary:
[reference generation] {creative goal using Subject 1}
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - {exact attributes}
<Video 1> (camera and pacing structure): weak_reference - {relationship}
detailed_description:
{style sentence}
[Shot 1] <Subject 1> ...
overall_soundscape:
...
non_diegetic_music:
...tracking shot + medium amplitude + slow speed(S1) says <d>[Chinese] 原句</d>instrumentation + tempo + dynamics; or N/A04 / REFERENCE ROLES
清楚的职责映射比增加更多参考素材更有用;必须保持的细节要在文字中再写一次。
<Subject 1>身份:脸、发型、服装
<Subject 2>来自 Picture 2 的产品:几何、材质、颜色
<Picture 3>具体首帧、关键帧或尾帧锚点
<Video 1>动作:步行节奏与镜头路径
<Audio 1>声音:仅参考音色
保持条件
In Video 1, change only the original package to the exact product from Picture 1.
Preserve the actor identity, hands, timing, camera path, background and lighting.
Keep the product geometry, color and logo from Picture 1 unchanged.PROMPT RECIPES
这些是原创教学模板,不代表必然复现;每次只调整一个变量。
10s premium product film. A matte-red speaker in clean negative space. Slow push-in as rim light traces the exact silhouette; cut once to a macro texture shot, then settle into a stable hero frame. One soft dial click, restrained room tone, minimal electronic percussion.
使用这个 PromptPicture 1 anchors 0.00s and Picture 2 anchors 8.00s. Use one continuous shot. The hand pulls the ribbon and the paper unfolds in observable stages. Every moving element gradually settles into the exact composition, lighting and hand position of Picture 2.
使用这个 PromptIn Video 1, change only the package to the exact product from Picture 1. Preserve actor identity, hands, timing, camera path, background, lighting and all other motion. Keep product geometry, color and logo unchanged.
使用这个 Prompt05 / TROUBLESHOOTING
每次重试只改变一个控制变量,才能判断哪条指令真正改善了结果。
| 症状 | 可能原因 | 下一步修改 |
|---|---|---|
| 人物或产品漂移 | 参考职责不明确,或同时承担风格和动作。 | 用一张干净图片只负责身份,并列出必须保持的具体特征。 |
| 镜头运动冲突 | 视频参考和文字描述了不同镜头路径。 | 让参考视频负责路径,文字只补充速度或幅度。 |
| 对白过快或声画错位 | 台词过长、speaker 变化,或没有绑定时间节点。 | 缩短台词、保持同一 speaker ID,并放进对应镜头。 |
| 首尾帧突然跳变 | 两帧的机位、尺度差异过大,或独立运动对象太多。 | 从首帧编辑出尾帧,只改变一个可达状态,并描述一条连续路径。 |
06 / CONTRACT BOUNDARIES
所有时长、分辨率、参数和 model ID 都必须标注具体由哪个平台提供。
官方 API 时长
4–15s
官方分辨率
768P / 2K
首尾帧比例
adaptive
文生视频比例
21:9 → 9:16
请求 schema、model string 与能力边界以官方文档为准。
官方 API 指南本站可用模式、控件、价格与限制以当前工作台显示为准。
打开工作台按 T2V、首尾帧或参考生成选择对应 workflow 与权重。
ComfyUI 指南FAQ
模式、参考素材、对白和平台合同都应该在风格润色前确定。
纯文字创意用 T2VA;首帧决定构图用 I2VA;首尾都重要用 FL2VA;预设结尾用 L2VA;需要图片、视频或音频控制特定属性时用 Ref2VA。
不要假设所有 Provider 都有该字段。改写成正向保持条件,例如“镜头保持固定”“保留完全一致的脸和服装”“嘴唇保持闭合”。
保持 speaker ID 一致,标注对白语言,让台词长度匹配时长,并把对白、环境声、物理音效和观众可听配乐分开。
不一样。权重、输入模式、时长、分辨率和参数都可能不同,必须以当前使用平台的合同为准。
来源与核验
内容最后核验:2026 年 8 月 11 日。页面模板是依据所链接结构原创编写的教学材料。