Seedance 2.5 is here! On annual plans, get flat 20% discount.

MiniMax H3

Every input. One take.

MiniMax H3 is an open-weights multimodal video model. Text, images, clips, and audio share one context and resolve into 5–15 seconds of 4K video with native stereo sound: score, dialogue, and foley mixed to picture.

On Chromatic you get text-to-video, first-and-last-frame, and reference-to-video (up to 9 images, 3 clips, 3 audio). Strong at stacked references, type rendering, and precise localized edits.

5-15sANY WHOLE SECOND LENGTH
4KPRODUCTION RESOLUTION ON CHROMATIC
A/VNATIVE STEREO IN THE SAME PASS
12REFS MAX: 9 STILLS, 3 CLIPS, 3 AUDIO

What H3 is built for.

Four strengths that change how you brief a shot: stacked references, instruction edits, type rendering, and native stereo audio.

Reference to video

Pass up to 9 images, 3 video clips, and 3 audio tracks in one generation (12 files max). Identity, motion, style, voice, and edit rhythm in a single brief.

Video editing

Upload a keeper and name the change in plain language. Swap a product, rewrite a sign, relight a scene, or replace a line while the rest of the frame stays put.

Type rendering

Titles, credits, packaging, and UI copy stay legible in motion when you quote the line, name the face, and say how it resolves.

Native audio

Every take returns stereo score, dialogue, foley, and room tone timed to picture. Direct the mix, or clone a voice from a reference recording.

What creators are building.

Brand films, prompt-only style, first-last motion, stacked refs, localized edits, physics with native audio.

Type rendering. Sci-fi mystery teaser. “THE STARS WERE LISTENING” resolves from soft blur into sharp condensed type. Face, lock, and color named in the prompt.

Text to video. Hand-drawn kitchen creature. Prompt-only: live-action dusk kitchen plus luminous animation, no reference stills.

First & last frame. Animated gallery poster. Composition already decided. H3 fills the motion between the still and the type-on.

Reference to video. Vintage binocular brand film. Four keyframes, each with a job (mood, scan, passerby, mark), plus red type on the rack-focus.

Instruction-based edits. Product, sign, and line swap. Coke in the bag, HUHUI on the storefront, new dialogue. The rest of the shot stays put.

Physics and native audio. Claymation lava canyon leap. Slow-motion fox, camera racing under the belly, clay body and chasm depth in one pass.

H3 vs Seedance.

H3 takes character refs Seedance often refuses. For a longer film you need reference-to-video; without a face in the brief, shots will not stay consistent. Type holds style. Timestamps land.

CREDITS

MINIMAX H3

4× less cost

SEEDANCE

Higher cost

FACES

Works smoothly with characters and faces. Seedance often denies the still, so you cannot lock identity in reference-to-video across a longer film.

MINIMAX H3

Works with faces

SEEDANCE

Often denies

ON-SCREEN TEXT

H3 holds the lettering and the style. Seedance sometimes blurs the type or does not pick up the style.

MINIMAX H3

Holds style

SEEDANCE

Blurs or misses

CONTROLS

H3 follows timed shot lists. Seedance sometimes drifts off the timestamp.

MINIMAX H3

Hits timestamps

SEEDANCE

Can drift

TIME

MINIMAX H3

8 min avg

SEEDANCE

4 min

Practical notes.

How to pick a mode, cite references, and structure a prompt that H3 can follow on Chromatic.

Prompting MiniMax H3: Key Techniques

Text to video

When you can fully describe the shot, write it like a director: world, action, and camera language H3 can execute. No reference stills required.

WORLD + SUBJECT

Name the set, talent, wardrobe, and hero prop so H3 invents a complete frame instead of a generic fashion void.

White infinity studio. Couture dress with butterfly-wing embroidery. Surreal red-fruit tree.

Instruction edits

Pass a clip, then list substitutions in plain language. Pair each change with what must stay stable so H3 edits instead of regenerating the shot.

CITE THE BASE CLIP

Name the reference so every later swap has a frame to edit against.

In the reference video…

Type rendering

Direct titles like a shot: exact copy, face, color, and how the letters resolve. Vague “add a title” is how you get mush.

BUILD THE FRAME FIRST

Give the title a place to live: scale, subject, and light so the type sits in depth instead of floating as a sticker.

Ultra-wide sci-fi corridor. A solitary figure with a flashlight faces a colossal circular gateway.

On Chromatic, you type a slash.

Chromatic understands MiniMax prompts. Paste what you have, select what matters, run the skill.

01Drop what you haveProduct shots, a face, a script, a rough note. Anywhere on the canvas.
02Select the pieces that matterThe skill only reads what you highlight.
/h3-
/h3-text-to-video
/h3-reference
03Type /h3- and goChromatic understands MiniMax prompts. Pick the skill and go.

MiniMax H3, on Chromatic.

Open the canvas. Light helps brief the shot. H3 is one of 50+ models on the platform.