Generate AI videos with MiniMax H3 Max — the fal.ai speed tier of MiniMax H3. Text, image, and reference to video with native stereo audio, right in your browser.
MiniMax H3 Max · fal.ai speed tier
H3 Max keeps H3’s prompt adherence, native stereo audio, and reference-guided consistency — and renders a 5-second 768p clip in under 3 seconds. Pick an input mode and generate right on this page.
H3 Max Showcase
Official showcase samples from the fal.ai MiniMax H3 Max model pages — text, image, and reference-driven clips with native audio, rendered in seconds. Turn the sound on, then try the model yourself above.
Demo 01
Demo 02
Demo 03
Demo 04
Generated with the MiniMax H3 Max model. Make your own →
Model overview
MiniMax H3 Max is a post-trained, speed-focused tier of the MiniMax H3 video model, hosted on fal.ai. It generates 5–15 second clips with native stereo audio and renders them faster than real time, so you can iterate on prompts the way you iterate on text.
| Model | MiniMax H3 Max (fal.ai post-trained tier of MiniMax H3) |
|---|---|
| Input modes | Text-to-video, image-to-video, reference-to-video |
| Duration | 5–15 seconds per clip |
| Resolution | 480P or 768P (H3 Max resolution tiers) |
| Audio | Native stereo audio on every clip |
| Speed | Community benchmarks: ~4.7s per text-to-video clip, ~6.4s per image-to-video |
Want the full model deep-dive with benchmarks and per-second pricing? Read the MiniMax H3 Max guide, or grab copy-paste prompts from our MiniMax H3 prompt guide.

Type a prompt with subject, camera, and audio cues — or upload an image to animate while keeping its subject and style.
Pick duration (5–15s) and resolution (480P/768P), hit generate, and get a finished clip with native audio in seconds.
Generate up to 15 seconds of 2K video with native stereo audio, powered by the MiniMax H3 model. Free to start — see the credit cost before you generate.