The MiniMax H3 model
MiniMax H3: Specs, Open Weights, Resolution, and API Cost
MiniMax H3 is an open general-purpose multimodal video model with 33B parameters and open weights, released on July 31, 2026. It generates up to 15 seconds of 2K video at 24fps with native stereo audio, accepts text, images, video, and audio as reference inputs, and ranks #1 on the Artificial Analysis video-editing leaderboard. This page is the complete technical reference: spec table, architecture, benchmarks, open-weight download options, the resolution ladder, and per-second API pricing.
MiniMax H3 quick specs
| Model type | Open general-purpose omni-modal video model |
| Parameters | 33B, with open weights published by MiniMax |
| Release date | July 31, 2026 |
| Max clip length | Up to 15 seconds |
| Max resolution | 2K native at 24fps (up to 4K on fal.ai) |
| Audio | Native stereo synchronized audio — speech, SFX, and ambience generated with the video |
| Input modes | Text-to-video, image-to-video, first/last-frame keyframes, reference-guided generation (Ref2VA), video-to-video editing |
| Reference inputs | Text, images, video, and audio mixed in one context — up to 9 reference inputs |
| Camera control | Pan, zoom, tilt, and roll, with timecode shot timing such as [0s-3s] |
| Prompt length | Up to 7,000 characters |
| Weights | Open-weight — downloadable from Hugging Face for local inference and fine-tuning |
| Rankings | #1 on the Artificial Analysis video-editing leaderboard and the Arena image-to-video board |
Architecture: an omni-modal video model
Most video models take one prompt and return one video. MiniMax H3 is built as a general-purpose omni-modal model: a single generation can condition on text, images, video clips, and audio at the same time — up to 9 mixed reference inputs in one context. That single design decision is what enables the workflows creators associate with the model: animating a product photo while holding the label pixel-perfect, or keeping the same spokesperson across every shot of a campaign.
The other half of the architecture story is audio. H3 generates the soundtrack natively with the visuals — dialogue, sound effects, and ambience in stereo, synchronized across the full clip. You direct it in the prompt itself, and the model treats sound as a first-class output rather than a post-processing step.
At 33B parameters with open weights, the model is large but not out of reach for the open-source ecosystem — which is exactly why MiniMax published it that way (see the open-weight section below).
Benchmarks and how the community rates it
Since release, MiniMax H3 has held #1 on the Artificial Analysis video-editing leaderboard and the Arena image-to-video board. Community reception centers on three strengths: image fidelity, editing control, and character consistency. In practical terms:
- Image fidelity: image-to-video outputs preserve the subject, style, and composition of the source photo instead of re-imagining them.
- Editing control: video-to-video edits respect shot boundaries, camera motion, and timing while swapping styles, characters, or products.
- Character consistency: reference-guided generation (Ref2VA) locks identity across shots — the same face, outfit, and product from clip to clip.
Honest limitations exist: occasional mouth and tongue artifacts on faces (worst on extreme close-ups), imperfect physics on fast motion, and slow rendering when you run the open weights locally. Every model in this class has trade-offs; H3's are concentrated in areas most edits can route around.
Open weights: Hugging Face, GitHub, and local deployment
Is MiniMax H3 open source? The weights are open: MiniMax published the 33B model on Hugging Face at release, so it can be downloaded, run locally, and fine-tuned without going through a proprietary API. For creators searching for the model on Hugging Face or GitHub: the official weights live under the MiniMax organization on Hugging Face, while GitHub hosts the community ecosystem around the model — local-inference helpers, quantized builds, and fine-tuning tooling maintained by third parties.
Running it yourself is the trade-off. Local rendering of a 33B video model on consumer hardware is currently slow — one clip can take a long stretch of GPU time, which breaks iterative workflows. Most creators pick one of two paths instead: a hosted API endpoint (fal.ai, per-second pricing below) or a browser generator like H3 Video, which runs MiniMax H3 in the cloud with nothing to install.
Resolution guide: 480p to 4K (and where H3 Max fits)
What resolution does MiniMax H3 support? Standard H3 generates up to 2K natively at 24fps, with 4K available on fal.ai. The H3 Maxvariant — post-trained by fal.ai for speed — tops out at 768p and hands 2K/4K and native-audio workloads back to standard H3. In practice the ladder looks like this:
| Tier | Availability | Typical use |
|---|---|---|
| 480p | Standard H3 and H3 Max | Fast drafts and volume production at the lowest per-second cost |
| 768p | Standard H3 and H3 Max (H3 Max max) | Social clips and iteration; H3 Max renders 5s at 768p in under 3 seconds |
| 2K native | Standard H3 | Deliverable-grade output at 24fps — the tier H3 is known for |
| 4K | Standard H3 on fal.ai | Master-grade renders for post-production pipelines |
A common workflow: prototype at 480p/768p on H3 Max for speed, then re-render the winning shot at 2K on standard H3. You get fast iteration and a high-resolution master without paying 2K prices on every attempt.
MiniMax H3 API cost
How much does the MiniMax H3 API cost? On fal.ai, MiniMax H3 is priced per second of generated video. Image-to-video on H3 Max is the cheapest entry point at $0.0125/s — a 15-second animated clip lands around $0.19 — while standard H3 at 2K and 4K commands the premium tiers:
| Task | 480p | 768p | 2K | 4K |
|---|---|---|---|---|
| MiniMax H3 — text-to-video | $0.05 / sec | $0.06 / sec | $0.13 / sec | $0.16 / sec |
| MiniMax H3 Max — text-to-video | $0.05 / sec | $0.08 / sec | — | — |
| MiniMax H3 Max — image-to-video | $0.0125 / sec | $0.02 / sec | — | — |
fal.ai also offers a free tier of up to 5 videos per day on the H3 Max model page. On H3 Video, generation is credit-based instead: you buy credit packs (no forced subscription), and the playground shows the exact credit cost before you generate — so the per-clip price is never a surprise.
The speed difference matters as much as the price: H3 Max renders a 5-second 768p clip in under 3 seconds — roughly 35x faster than the official MiniMax H3 endpoint — while standard H3 prioritizes fidelity over latency. For a deeper comparison, see the MiniMax H3 Max guide.
Capabilities you can use today
Four input modes
- Text-to-video: a single prompt becomes a full scene with motion, camera work, and sound — prompts run up to 7,000 characters, long enough for a complete shot list.
- Image-to-video: animate a photo while preserving your subject, style, and composition.
- Reference-guided generation (Ref2VA): attach up to 9 references — images, video, audio — to lock characters, products, and styles across shots.
- Video-to-video editing: restyle or extend existing footage with precise multimodal control, the capability behind its #1 editing ranking.
Cinematic camera control
H3 reads cinematography vocabulary directly: pan, zoom, tilt, and roll, plus timecode shot timing like [0s-3s] for a push-in followed by [3s-8s] for a slow orbit — so you can storyboard camera moves inside a single 15-second clip. For prompt technique, see the MiniMax H3 prompt guide with 25+ copy-paste examples.
Generate MiniMax H3 videos on H3 Video
H3 Video is an independent generator powered by the MiniMax H3 model: pick an input mode, write your prompt or upload media, see the credit cost, and generate — nothing to install, no local GPU, native audio on every clip. Start in the playground, or jump straight to the MiniMax H3 Max and H3 Max Turbo fast generators.
Create your first H3 AI video
Generate up to 15 seconds of 2K video with native stereo audio, powered by the MiniMax H3 model. See the credit cost before you generate — free to start.
Open the H3 Video generatorMiniMax H3 FAQ
Is MiniMax H3 open source?
MiniMax H3 ships with open weights: the 33B-parameter model is published on Hugging Face, so anyone can download it, run it locally, and fine-tune it. Local rendering on consumer hardware is currently slow, which is why most creators use a hosted endpoint or a browser generator like H3 Video instead.
Where can I download MiniMax H3 — Hugging Face or GitHub?
The official weights live on Hugging Face under the MiniMax organization — search "MiniMax H3" there. GitHub is where the community collects local-inference helpers, quantized builds, and fine-tuning tooling around the model.
What resolution does MiniMax H3 support?
Standard MiniMax H3 generates up to 2K natively at 24fps, with 4K available on fal.ai. The H3 Max variant is tuned for speed and tops out at 768p; 480p and 768p are the budget tiers for standard H3 as well.
How much does the MiniMax H3 API cost?
On fal.ai, standard MiniMax H3 text-to-video costs $0.05/s at 480p, $0.06/s at 768p, $0.13/s at 2K, and $0.16/s at 4K. H3 Max ranges from $0.0125/s (image-to-video, 480p) to $0.08/s (text-to-video, 768p). On H3 Video you buy credit packs and see the exact cost before every generation.
Does MiniMax H3 generate audio?
Yes — native stereo synchronized audio covering speech, sound effects, and ambience is generated together with the video, across the full clip duration.
How long can a MiniMax H3 video be?
Up to 15 seconds at up to 2K resolution and 24fps, with native audio across the full duration.
What is the difference between MiniMax H3 and MiniMax H3 Max?
H3 Max is a variant post-trained by fal.ai for speed: 5-second 768p clips in under 3 seconds, at the cost of the 2K/4K tiers and the focus on native audio. Use H3 Max to iterate fast and standard H3 for the final master.
Is H3 Video the official MiniMax H3 site?
No. H3 Video is an independent platform powered by the MiniMax H3 model. H3 Video is not affiliated with, endorsed by, or sponsored by MiniMax or Hailuo AI.
This page is an independent technical reference written by the H3 Video team and is not affiliated with, sponsored by, or endorsed by MiniMax, Hailuo AI, or fal.ai. MiniMax, Hailuo, and fal are trademarks of their respective owners. Specs, rankings, and per-second pricing reflect vendor-published information (MiniMax's H3 announcement and fal.ai's model pages) as of September 2026 and may change without notice.