Camera & motion
01Direct the move, not just the scene
Describe the subject, camera path, pace, and sound in one prompt. H3 Max follows ordered shot instructions, helping tracking shots, push-ins, and continuous takes hold the intended visual beat.
Text · Image · Keyframes · Audio
Turn text, keyframes, or creative references into a 5–15 second video with synchronized audio. Direct camera motion, character continuity, story beats, and sound in one brief.
Prompt
0/2000
Fast direction, synchronized sound.
480p / 768p · 5–15 seconds
What it is
MiniMax H3 Max is a post-trained variant of the open-weight MiniMax H3 video model, optimized for prompt adherence, visual quality, and fast inference. On Waver AI, it creates video from text, first and last frames, or mixed references—and generates synchronized audio in the same pass.
Independent evaluation
At launch, fal reported leading image-to-video results on Design Arena and Artificial Analysis. Rankings change, so these are presented as dated launch snapshots.


Key features
Combine shot direction, visual consistency, and sound in one prompt-led workflow.
Camera & motion
01Describe the subject, camera path, pace, and sound in one prompt. H3 Max follows ordered shot instructions, helping tracking shots, push-ins, and continuous takes hold the intended visual beat.
First & last frame
02Start with an opening image and optionally add a closing image. H3 Max builds the movement between those keyframes, giving image-to-video work a defined destination.
Character consistency
03Keep a person recognizable as lighting, framing, and location change. The model preserves core facial features, clothing, and proportions across a multi-shot sequence.
Native audio
04Write dialogue, ambience, music, and foley into the same brief as the visuals. H3 Max returns synchronized sound with the picture, reducing separate soundtrack work for quick concepts and social clips.
Art direction
05Name the palette, material, linework, and typography you want. H3 Max can carry a defined visual treatment across several beats for more coherent brand pieces and animated stories.
Prompt adherence
06Structured prompts can specify what happens first, what changes, and how the shot resolves. Post-training focuses on adherence and aesthetics so detailed briefs survive the move from words to motion.
Showcase
Explore cinematic motion, stylized scenes, dialogue, character continuity, and synchronized audio.
How it works
Start from a text prompt, an opening image, or a first-and-last-frame pair when the shot needs a specific ending.
Describe the subject, action, camera movement, pacing, lighting, and audio in the order they should happen.
Choose 5–15 seconds, select 480p or 768p, confirm the displayed credit cost, and generate the clip.
Use cases
Turn a product brief into a compact clip with camera direction, readable visual beats, and synchronized sound. Generate quick concept options before a campaign enters final production.
Test a scene, transition, or camera move before committing to a full shoot. Fast 768p output makes comparing directions easier while the story is still flexible.
Describe a speaker, line, framing, and background sound in one prompt. Native audio and lip sync make short character moments easier to review as complete ideas.
Carry a deliberate palette and material language across several shots for animated posters, product reveals, and branded scenes that need consistent art direction.
Credit pricing
480p
1 credit / second
Best for quick drafts, shot tests, and rapid iteration.
768p
2 credits / second
Sharper output for review, presentation, and social delivery.
Reference mode adds 1 credit per input image. Reference-video duration is charged at the selected output rate.
Model choice
Explore the complete MiniMax video model family, or compare credit bundles on the Waver AI pricing page.
FAQ
MiniMax H3 Max is a post-trained version of the open-weight MiniMax H3 base model, tuned for stronger prompt adherence, aesthetics, and fast 768p inference. Waver AI supports text, image, and reference-led video workflows.
It can generate video from text, animate a starting image with an optional final keyframe, or follow image, video, and audio references. It supports directed camera motion, consistent characters and styles, and synchronized audio cues.
A single generation can run from 5 to 15 seconds. Short clips suit quick shot tests, while 15 seconds can carry several ordered beats or a compact multi-shot sequence.
MiniMax H3 Max supports 480p and 768p output and is optimized around fast 768p creation. Standard MiniMax H3 is a better fit when a workflow specifically needs 2K output.
Yes. It generates synchronized audio with the picture. Prompts can describe dialogue, foley, room tone, ambience, or music so sound arrives in the same pass.
No. H3 Max prioritizes speed, prompt adherence, and aesthetics at 480p or 768p. Standard H3 supports 2K and a broader production-oriented multimodal workflow.
Yes. Upload a starting image and optionally an ending image. The output follows the source image ratio, while text-to-video offers landscape, square, and portrait ratios.
480p output uses 1 credit per second and 768p uses 2 credits per second. Reference mode also adds 1 credit per reference image and charges reference-video seconds at the selected output rate.
Create on Waver AI
Write a shot brief or upload a starting image, choose your settings, and generate a synchronized MiniMax H3 Max video.
Create with MiniMax H3 Max