T2VA
Text to Video
Text prompt
Open-ended concepts, ads and cinematic scenes.
Describe the initial composition, one clear action, camera path and sound.
Open this workflowChoose the right generation mode, build a prompt with the official H3 structure, assign reference roles, and fix common failures without guessing.
01 / MODE PICKER
Start here. Picking the wrong mode creates conflicts before the model reads a single camera instruction.
T2VA
Text prompt
Open-ended concepts, ads and cinematic scenes.
Describe the initial composition, one clear action, camera path and sound.
Open this workflowI2VA
First image + prompt
Identity, product shape and composition control.
The image owns the static frame; the prompt should explain what moves next.
Open this workflowFL2VA
First image + last image + prompt
Transformations and exact end-state convergence.
Use one achievable motion path; keep the two frames compositionally compatible.
Open this workflowL2VA
Last image + prompt
Logo reveals, product end cards and designed endings.
Describe a plausible earlier state and how every element settles into the supplied frame.
Open this workflowRef2VA
Images, video, audio + prompt
Identity, product, motion, camera or voice transfer.
Assign one explicit job to every asset and state what must be preserved.
Open this workflowCONTROL RULES
Use these rules before adding cinematic detail. More words do not fix a conflicted brief.
Input type changes what the prompt should control. Do not write one universal prompt for every workflow.
Short clips need one readable event. Add a new shot only when the visual information genuinely changes.
Separate identity, product, motion, camera and voice roles so the references do not compete.
Replace vague negatives with explicit states: static camera, closed lips, exact face, unchanged wardrobe.
02 / PROMPT BUILDER
The builder changes its alignment instruction, reference labels and API content roles for the selected H3 mode.
Generation mode
10s. A matte-red portable speaker remains geometrically consistent. A rim light traces the silhouette, then the speaker settles into a clean hero frame. Minimal dark studio, soft haze, precise reflections and generous negative space. The camera performs a Push In with small amplitude at slow speed. One soft dial click, restrained room tone and a synchronized low-frequency pulse. Minimal electronic percussion, moderate tempo, ending cleanly on the hero frame.The output updates as you type. Keep structural prompts in English; use the target language for dialogue and on-screen text.
462/700003 / OFFICIAL ANATOMY
Use the field names as a thinking system: visuals and shots first, diegetic sound second, audience-only music last.
T2VA / I2VA / FL2VA / L2VA
{I2VA / FL2VA / L2VA alignment instruction when applicable}
integrated_multimodal_description:
[Shot 1] {initial composition}. {subject action}.
[Shot 2] At 00:04.000, {new visual information}.
overall_soundscape:
{ambience + physical sounds}
non_diegetic_music:
{instrumentation + tempo + dynamics}; or N/ARef2VA
subject_definitions:
<Subject 1> is the product shown in <Picture 1>; preserve its geometry, material, color and logo.
<Video 1> is the camera-path and pacing reference.
summary:
[reference generation] {creative goal using Subject 1}
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - {exact attributes}
<Video 1> (camera and pacing structure): weak_reference - {relationship}
detailed_description:
{style sentence}
[Shot 1] <Subject 1> ...
overall_soundscape:
...
non_diegetic_music:
...tracking shot + medium amplitude + slow speed(S1) says <d>[Chinese] 原句</d>instrumentation + tempo + dynamics; or N/A04 / REFERENCE ROLES
A clean role map is more useful than adding more references. Repeat the must-preserve details in text.
<Subject 1>Identity: face, hair, wardrobe
<Subject 2>Product from Picture 2: geometry, material, color
<Picture 3>Concrete first, key or last-frame anchor
<Video 1>Motion: walking rhythm and camera path
<Audio 1>Voice: timbre only
Preservation contract
In Video 1, change only the original package to the exact product from Picture 1.
Preserve the actor identity, hands, timing, camera path, background and lighting.
Keep the product geometry, color and logo from Picture 1 unchanged.PROMPT RECIPES
These are original teaching templates, not claims of guaranteed output. Adapt one variable at a time.
10s premium product film. A matte-red speaker in clean negative space. Slow push-in as rim light traces the exact silhouette; cut once to a macro texture shot, then settle into a stable hero frame. One soft dial click, restrained room tone, minimal electronic percussion.
Use this promptPicture 1 anchors 0.00s and Picture 2 anchors 8.00s. Use one continuous shot. The hand pulls the ribbon and the paper unfolds in observable stages. Every moving element gradually settles into the exact composition, lighting and hand position of Picture 2.
Use this promptIn Video 1, change only the package to the exact product from Picture 1. Preserve actor identity, hands, timing, camera path, background, lighting and all other motion. Keep product geometry, color and logo unchanged.
Use this prompt05 / TROUBLESHOOTING
Change one control variable per retry so you can tell which instruction improved the result.
| Symptom | Likely cause | Change next |
|---|---|---|
| Face or product drifts | The reference role is implicit or shared with style and motion. | Assign identity to one clean image and list the exact features to preserve. |
| Camera motion conflicts | The video reference and text describe different camera paths. | Let the reference own the path; use text only for speed or amplitude. |
| Dialogue is rushed or mismatched | The line is too long, the speaker changes, or timing is not anchored. | Shorten the line, keep one speaker ID and place it inside the matching shot. |
| First/last frame jumps | The frames differ in camera angle, scale or too many independent objects. | Create the last frame from the first, change one reachable state and describe one continuous path. |
06 / CONTRACT BOUNDARIES
Label every duration, resolution, parameter and model ID by the surface that actually provides it.
Official API duration
4–15s
Official resolution
768P / 2K
Frame-mode ratio
adaptive
T2V ratio
21:9 → 9:16
Use the official request schema, model string and capability documentation as the source of truth.
Official API guideUse only the modes, controls, prices and limits visible in the current workspace.
Open workspaceSelect the workflow and weights that match T2V, image-frame or reference generation.
ComfyUI guideFAQ
Mode, references, dialogue and platform contracts should be decided before style polish.
Use T2VA for text-only concepts, I2VA when a first frame owns the composition, FL2VA when both endpoints matter, L2VA for a designed ending, and Ref2VA when existing images, video or audio must control specific attributes.
Do not assume one exists across providers. Write positive preservation conditions instead: “camera holds a static shot,” “preserve the exact face and wardrobe,” or “lips remain closed.”
Keep speaker IDs stable, label dialogue language, keep lines short enough for the duration, and separate dialogue from ambience, physical sounds and audience-only music.
No. Model weights, input modes, duration, resolution and parameters can differ. Always read the contract for the surface you are using.
Sources & verification
Last content verification: August 11, 2026. Templates on this page are original teaching material based on the linked structures.
Pick the workflow that matches your source material, build the prompt above, then carry it into the corresponding generator.