MiniMax H3 Max, post-trained by fal for speed
MiniMax H3 Max renders a 5-second clip in under 3 seconds
MiniMax H3 Max is the speed-tuned build of MiniMax H3, post-trained by fal. In fal's September 2026 benchmark it rendered a 5-second 768P clip in about 2.5 seconds and a 15-second 1080P clip in under 18 seconds, roughly 15x faster than MiniMax's own H3 inference, and as of September 2026 it ranks first for image-to-video on the Artificial Analysis and Design Arena leaderboards. Create from text, animate first and last frames, combine multimodal references, or direct keyframed camera movement, then iterate on 5–15 second clips at 480P, 768P, or 1080P while ideas are still fresh.
Four generation modes
Start MiniMax H3 Max from an idea, a composition, a set of references, or a camera path. Its fast generation makes it practical to test more directions while each mode keeps a different creative variable under control.
Text to video
Describe the subject, action, environment, lens behavior, lighting, and mood, then let text-to-video generation return a complete scene in seconds, so a rewritten prompt can be checked right away.
Image to video
Use one image as the exact opening frame and optionally add a last frame when the shot needs a defined visual destination.
Reference to video
Combine up to 12 images, video clips, and audio files to guide character identity, style, movement, timing, and atmosphere.
Camera Controls
Keep the reference scene frozen while ordered keyframes control camera time, orbit angle, elevation, and relative distance.
A practical fast-video workflow
- 1Choose the mode that matches what must stay predictable: the concept, opening composition, reference identity, or camera path.
- 2Write one readable action and camera direction, then select duration, resolution, aspect ratio where available, and prompt-expansion quality for stronger prompt adherence.
- 3Generate a first version quickly, review continuity, motion, composition, and lighting, then change one instruction or reference at a time for the next fast iteration.
Create more directions in less time
- Help marketing teams turn product stills into polished reveals and camera-controlled campaign concepts on a rapid production cycle.
- Give creators and studios a fast way to animate portraits, illustrations, landscapes, and storyboards while preserving the opening composition.
- Let creative teams test consistent characters, visual styles, motion, and audio-led references across more ideas before choosing a final direction.
Turn generation speed into a repeatable creative process
MiniMax H3 Max is most useful when speed supports clear decisions instead of producing random variations. Begin with a narrow question for each draft: does the subject move correctly, does the camera follow the intended path, or do the references preserve the right identity and mood? Fast turnaround lets you answer one question, keep what works, and move to the next without losing the original creative direction.
Test motion before raising resolution
Use a short 480P draft to check whether the main action reads clearly and finishes inside the selected duration. Look for unwanted changes in the subject, background, framing, or transition. Once the movement and timing are convincing, repeat the established direction at 768P or 1080P for a more polished result instead of spending the higher-resolution pass on an unproven idea.
Write prompts around visible change
Structure the prompt in the order a viewer experiences the shot. Identify the subject and the details that must remain recognizable, describe one primary action, specify the camera response, then add lighting, atmosphere, pacing, and the desired ending. Concrete visual instructions improve prompt adherence more reliably than a long list of unrelated adjectives, especially when several references already provide style and identity.
Separate subject motion from camera motion
In regular text or image generation, state what the subject does and how the camera observes it as two distinct instructions. When the scene should remain completely still, switch to Camera Controls and build an ordered trajectory instead. Adjust normalized time, azimuth, elevation, and distance deliberately so an orbit, rise, pullback, or approach can be evaluated and repeated without asking the model to invent scene movement.
Review one variable per iteration
After each result, compare it with the original brief before changing anything. Check identity, composition, action, camera path, lighting, reference influence, and the ending separately. Keep successful settings and revise only the weakest variable in the next fast generation. This disciplined loop makes improvements easier to identify and prevents a useful prompt, frame pair, or camera trajectory from being lost among broad rewrites.
H3 or H3 Max: which model to use
Both models share the MiniMax H3 architecture and the same 5–15 second range, but they are tuned for different jobs. H3 Max is post-trained by fal for generation speed, prompt adherence, and keyframed camera control. The original MiniMax H3 keeps 2K output and larger multimodal reference briefs. Choose by what the shot needs, not by the name.
| Feature | MiniMax H3 | MiniMax H3 Max |
|---|---|---|
| Best for | 2K delivery and multimodal reference briefs | Fast iteration and keyframed camera moves |
| Generation speed | Standard MiniMax inference | 5-second 768P clip in about 2.5 seconds, roughly 15x faster |
| Resolution | 768P, 2K | 480P, 768P, 1080P |
| Duration | 5–15 seconds | 5–15 seconds |
| Camera control | Described in the prompt | Up to 12 ordered keyframes: time, azimuth, elevation, distance |
| References | Up to 9 images, 3 videos, and 3 audio clips | Up to 12 files across images, videos, and audio |
Speed figures are fal's own timings for the H3 Max endpoint, measured on September 8, 2026. Real wait time adds queue, prompt expansion, and encoding.
Need 2K output? Open MiniMax H3MiniMax H3 Max frequently asked questions
What is MiniMax H3 Max?
MiniMax H3 Max is an exceptionally fast AI video model with text, image, multimodal reference, and keyframe-based camera-control workflows, designed for prompt-faithful results and rapid creative iteration.
What duration and resolution can I select?
MiniMax H3 Max supports whole-second durations from 5 to 15 seconds at 480P, 768P, or 1080P.
How many reference files can I upload?
Reference mode accepts up to 12 files across images, videos, and audio. Refer to each asset in the prompt by its media type and order.
How do Camera Controls work?
Add up to 12 ordered keyframes. Each keyframe sets normalized time, horizontal azimuth, vertical elevation, and relative camera distance.