MiniMax H3 multimodal video generation
Build more controllable shots with MiniMax H3
MiniMax H3 is an open, general-purpose multimodal video model designed for creation and editing. It can combine a prompt with image, video, and audio references, preserve visual direction through first and last frames, and deliver 2K clips with enough duration for a complete beat rather than a fleeting loop.
One multimodal model, many ways to create
Start with text, a first frame, a first-and-last-frame pair, or a set of visual and audio references. MiniMax H3 can generate a new shot, extend a visual idea, or edit existing video while keeping the intended subject, movement, and atmosphere in view.
Read the official MiniMax H3 documentationDirect every shot with richer references
Use first and last frames to define a visual arc, or combine up to 9 images, 3 videos, and 3 audio clips within a 12-file reference set. MiniMax H3 reads those signals together, so character identity, camera language, rhythm, and mood can reinforce one another.
Generation modes for every starting point
MiniMax H3 adapts to the material you already have, from a written concept to a complete clip that needs a targeted edit.
Generate from a text prompt when you want to explore a new scene, action, camera move, or visual style from scratch.
Use a first frame to anchor composition, add a last frame to guide the ending, or combine both for a more intentional transition.
Bring image, video, and audio references into one brief, or edit an existing clip while preserving the creative direction that matters.
A practical MiniMax H3 workflow
A focused brief helps the model turn multimodal context into coherent motion. Start with the strongest reference, then add only the signals that improve the shot.
- 1Choose text-to-video, frame-guided generation, reference generation, or video editing based on the asset you want to create.
- 2Describe one clear subject action, the camera movement, lighting, atmosphere, and how the shot should settle.
- 3Set a 5–15-second duration, generate the shot, then refine one variable at a time so each iteration has a clear purpose.
Where MiniMax H3 fits
- Animate product stills into short reveals, rotating details, and landing-page motion.
- Turn character art or portraits into controlled micro-scenes with a clear camera move.
- Prototype social clips and campaign ideas before investing in longer production.
How to write a useful MiniMax H3 prompt
Treat the prompt as a compact director's brief. References establish identity and texture; your words should make motion, camera, timing, and the ending unambiguous.
Subject motion
Give the main subject one readable action that can develop and resolve within the selected 5–15-second shot.
Camera direction
State a specific move such as a slow push-in, orbit, tracking shot, tilt, or locked camera.
Visual continuity
Name the lighting, atmosphere, background motion, and desired ending without replacing the reference image's identity.
Build a focused MiniMax H3 reference set
A larger reference set is not automatically better. Choose the smallest combination that establishes composition, subject identity, motion, sound, and the ending, then use the prompt to resolve any remaining ambiguity.
Start with the minimum input
Use text when exploring a concept from scratch. Add a first frame when composition or character appearance matters, and add a last frame only when the final pose, framing, or transition must land in a specific place.
Give every reference one job
Before uploading, define the purpose of each image, video, or audio clip in your brief. Keep the set within the supported limits—up to 9 images, 3 videos, 3 audio clips, and 12 files total—and remember that audio cannot be the only input.
Prepare clear source material
Crop still images around the subject you want to preserve, trim reference video to the movement that matters, and select the audio passage that establishes the intended rhythm or atmosphere. Avoid conflicting overlays or unrelated subjects that compete for attention.
Review the complete shot
Check whether the subject remains recognizable, the action finishes within the selected duration, and the camera move follows the prompt. If one part drifts, change that instruction or reference alone before generating again so you can identify which signal improved the result.
Use one evaluation goal per iteration
Separate technical success from creative judgment. First confirm that the requested duration, framing, and reference roles were followed. Then review identity, motion path, timing, lighting, background behavior, sound, and the ending against your brief. Save the prompt and reference combination that produced the strongest result before testing another variation, so a useful setup is never lost between iterations.
MiniMax H3 frequently asked questions
What is MiniMax H3?
MiniMax H3 is MiniMax's open, general-purpose multimodal video model. It is built for video generation and editing from text, images, video, and audio context.
What can I use as a reference?
You can guide a shot with a first frame, a last frame, or multimodal references. Reference mode accepts up to 9 images, 3 videos, and 3 audio clips, with no more than 12 files in total; audio cannot be the only input.
What resolution and duration does MiniMax H3 support?
The V2 API supports MiniMax H3 at 2K with whole-second durations from 5 to 15 seconds. It accepts text; first and/or last-frame images; or reference images, videos, and audio. Reference mode supports up to 9 images, 3 video clips, and 3 audio clips, capped at 12 files total; audio cannot be sent alone.
How does MiniMax H3 fit into the Hailuo model family?
MiniMax H3 is the model identifier used in MiniMax's current documentation. In this workspace, it appears as a coming-soon option in the Hailuo family until API access is connected.