03:00:0050% OFFClaim
Mode-aware guide + live builder

MiniMax H3Prompt Builder

Choose the right generation mode, build a prompt with the official H3 structure, assign reference roles, and fix common failures without guessing.

01 / MODE PICKER

The input decides what your prompt must do

Start here. Picking the wrong mode creates conflicts before the model reads a single camera instruction.

T2VA

Text to Video

Text prompt

Open-ended concepts, ads and cinematic scenes.

Describe the initial composition, one clear action, camera path and sound.

Open this workflow

I2VA

First Frame

First image + prompt

Identity, product shape and composition control.

The image owns the static frame; the prompt should explain what moves next.

Open this workflow

FL2VA

First + Last Frame

First image + last image + prompt

Transformations and exact end-state convergence.

Use one achievable motion path; keep the two frames compositionally compatible.

Open this workflow

L2VA

Last Frame

Last image + prompt

Logo reveals, product end cards and designed endings.

Describe a plausible earlier state and how every element settles into the supplied frame.

Open this workflow

Ref2VA

Reference to Video

Images, video, audio + prompt

Identity, product, motion, camera or voice transfer.

Assign one explicit job to every asset and state what must be preserved.

Open this workflow

CONTROL RULES

Four rules prevent most prompt failures

Use these rules before adding cinematic detail. More words do not fix a conflicted brief.

01

Choose the mode first

Input type changes what the prompt should control. Do not write one universal prompt for every workflow.

02

Match action to duration

Short clips need one readable event. Add a new shot only when the visual information genuinely changes.

03

Give each reference one job

Separate identity, product, motion, camera and voice roles so the references do not compete.

04

Preserve with positive language

Replace vague negatives with explicit states: static camera, closed lips, exact face, unchanged wardrobe.

02 / PROMPT BUILDER

Build once, then compare quick, H3 structure and API outputs

The builder changes its alignment instruction, reference labels and API content roles for the selected H3 mode.

Generation mode

10s. A matte-red portable speaker remains geometrically consistent. A rim light traces the silhouette, then the speaker settles into a clean hero frame. Minimal dark studio, soft haze, precise reflections and generous negative space. The camera performs a Push In with small amplitude at slow speed. One soft dial click, restrained room tone and a synchronized low-frequency pulse. Minimal electronic percussion, moderate tempo, ending cleanly on the hero frame.

The output updates as you type. Keep structural prompts in English; use the target language for dialogue and on-screen text.

462/7000
Use in generator

03 / OFFICIAL ANATOMY

Base and reference modes use different structures

Use the field names as a thinking system: visuals and shots first, diegetic sound second, audience-only music last.

T2VA / I2VA / FL2VA / L2VA

Base structure

{I2VA / FL2VA / L2VA alignment instruction when applicable}

integrated_multimodal_description:
[Shot 1] {initial composition}. {subject action}.
[Shot 2] At 00:04.000, {new visual information}.

overall_soundscape:
{ambience + physical sounds}

non_diegetic_music:
{instrumentation + tempo + dynamics}; or N/A

Ref2VA

Reference structure

subject_definitions:
<Subject 1> is the product shown in <Picture 1>; preserve its geometry, material, color and logo.
<Video 1> is the camera-path and pacing reference.

summary:
[reference generation] {creative goal using Subject 1}

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - {exact attributes}
<Video 1> (camera and pacing structure): weak_reference - {relationship}

detailed_description:
{style sentence}
[Shot 1] <Subject 1> ...

overall_soundscape:
...

non_diegetic_music:
...

Camera

tracking shot + medium amplitude + slow speed

Dialogue

(S1) says <d>[Chinese] 原句</d>

Music

instrumentation + tempo + dynamics; or N/A

04 / REFERENCE ROLES

Do not upload assets without assigning control

A clean role map is more useful than adding more references. Repeat the must-preserve details in text.

<Subject 1>

Identity: face, hair, wardrobe

<Subject 2>

Product from Picture 2: geometry, material, color

<Picture 3>

Concrete first, key or last-frame anchor

<Video 1>

Motion: walking rhythm and camera path

<Audio 1>

Voice: timbre only

Preservation contract

In Video 1, change only the original package to the exact product from Picture 1.

Preserve the actor identity, hands, timing, camera path, background and lighting.

Keep the product geometry, color and logo from Picture 1 unchanged.
Editing pattern: name the change once, then enumerate everything that must remain stable.

PROMPT RECIPES

Three reusable recipes, each with one clear job

These are original teaching templates, not claims of guaranteed output. Adapt one variable at a time.

T2VA

Premium product hero

10s premium product film. A matte-red speaker in clean negative space. Slow push-in as rim light traces the exact silhouette; cut once to a macro texture shot, then settle into a stable hero frame. One soft dial click, restrained room tone, minimal electronic percussion.

Use this prompt
FL2VA

Reach the designed end frame

Picture 1 anchors 0.00s and Picture 2 anchors 8.00s. Use one continuous shot. The hand pulls the ribbon and the paper unfolds in observable stages. Every moving element gradually settles into the exact composition, lighting and hand position of Picture 2.

Use this prompt
Ref2VA

Replace one product only

In Video 1, change only the package to the exact product from Picture 1. Preserve actor identity, hands, timing, camera path, background, lighting and all other motion. Keep product geometry, color and logo unchanged.

Use this prompt

05 / TROUBLESHOOTING

Fix the cause, not the adjective count

Change one control variable per retry so you can tell which instruction improved the result.

SymptomLikely causeChange next
Face or product driftsThe reference role is implicit or shared with style and motion.Assign identity to one clean image and list the exact features to preserve.
Camera motion conflictsThe video reference and text describe different camera paths.Let the reference own the path; use text only for speed or amplitude.
Dialogue is rushed or mismatchedThe line is too long, the speaker changes, or timing is not anchored.Shorten the line, keep one speaker ID and place it inside the matching shot.
First/last frame jumpsThe frames differ in camera angle, scale or too many independent objects.Create the last frame from the first, change one reachable state and describe one continuous path.

06 / CONTRACT BOUNDARIES

Official API, this studio and local workflows are not one contract

Label every duration, resolution, parameter and model ID by the surface that actually provides it.

Official API duration

4–15s

Official resolution

768P / 2K

Frame-mode ratio

adaptive

T2V ratio

21:9 → 9:16

MiniMax official API

Use the official request schema, model string and capability documentation as the source of truth.

Official API guide

MiniMax H3 AI Video Studio

Use only the modes, controls, prices and limits visible in the current workspace.

Open workspace

Local / ComfyUI

Select the workflow and weights that match T2V, image-frame or reference generation.

ComfyUI guide

FAQ

Questions that change the prompt

Mode, references, dialogue and platform contracts should be decided before style polish.

Which MiniMax H3 mode should I choose?

Use T2VA for text-only concepts, I2VA when a first frame owns the composition, FL2VA when both endpoints matter, L2VA for a designed ending, and Ref2VA when existing images, video or audio must control specific attributes.

Does MiniMax H3 have a universal negative prompt field?

Do not assume one exists across providers. Write positive preservation conditions instead: “camera holds a static shot,” “preserve the exact face and wardrobe,” or “lips remain closed.”

How should dialogue and sound be written?

Keep speaker IDs stable, label dialogue language, keep lines short enough for the duration, and separate dialogue from ambience, physical sounds and audience-only music.

Are hosted API and local ComfyUI limits the same?

No. Model weights, input modes, duration, resolution and parameters can differ. Always read the contract for the surface you are using.

Sources & verification

Use first-party facts, then label every provider difference

Last content verification: August 11, 2026. Templates on this page are original teaching material based on the linked structures.

Start with the mode, not a blank prompt box

Pick the workflow that matches your source material, build the prompt above, then carry it into the corresponding generator.