The MiniMax H3 model

MiniMax H3: Specs, Open Weights, Resolution, and API Cost

MiniMax H3 is an open general-purpose multimodal video model with 33B parameters and open weights, released on July 31, 2026. It generates up to 15 seconds of 2K video at 24fps with native stereo audio, accepts text, images, video, and audio as reference inputs, and ranks #1 on the Artificial Analysis video-editing leaderboard. This page is the complete technical reference: spec table, architecture, benchmarks, open-weight download options, the resolution ladder, and per-second API pricing.

MiniMax H3 quick specs

Model typeOpen general-purpose omni-modal video model
Parameters33B, with open weights published by MiniMax
Release dateJuly 31, 2026
Max clip lengthUp to 15 seconds
Max resolution2K native at 24fps (up to 4K on fal.ai)
AudioNative stereo synchronized audio — speech, SFX, and ambience generated with the video
Input modesText-to-video, image-to-video, first/last-frame keyframes, reference-guided generation (Ref2VA), video-to-video editing
Reference inputsText, images, video, and audio mixed in one context — up to 9 reference inputs
Camera controlPan, zoom, tilt, and roll, with timecode shot timing such as [0s-3s]
Prompt lengthUp to 7,000 characters
WeightsOpen-weight — downloadable from Hugging Face for local inference and fine-tuning
Rankings#1 on the Artificial Analysis video-editing leaderboard and the Arena image-to-video board

Architecture: an omni-modal video model

Most video models take one prompt and return one video. MiniMax H3 is built as a general-purpose omni-modal model: a single generation can condition on text, images, video clips, and audio at the same time — up to 9 mixed reference inputs in one context. That single design decision is what enables the workflows creators associate with the model: animating a product photo while holding the label pixel-perfect, or keeping the same spokesperson across every shot of a campaign.

The other half of the architecture story is audio. H3 generates the soundtrack natively with the visuals — dialogue, sound effects, and ambience in stereo, synchronized across the full clip. You direct it in the prompt itself, and the model treats sound as a first-class output rather than a post-processing step.

At 33B parameters with open weights, the model is large but not out of reach for the open-source ecosystem — which is exactly why MiniMax published it that way (see the open-weight section below).

Benchmarks and how the community rates it

Since release, MiniMax H3 has held #1 on the Artificial Analysis video-editing leaderboard and the Arena image-to-video board. Community reception centers on three strengths: image fidelity, editing control, and character consistency. In practical terms:

Honest limitations exist: occasional mouth and tongue artifacts on faces (worst on extreme close-ups), imperfect physics on fast motion, and slow rendering when you run the open weights locally. Every model in this class has trade-offs; H3's are concentrated in areas most edits can route around.

Open weights: Hugging Face, GitHub, and local deployment

Is MiniMax H3 open source? The weights are open: MiniMax published the 33B model on Hugging Face at release, so it can be downloaded, run locally, and fine-tuned without going through a proprietary API. For creators searching for the model on Hugging Face or GitHub: the official weights live under the MiniMax organization on Hugging Face, while GitHub hosts the community ecosystem around the model — local-inference helpers, quantized builds, and fine-tuning tooling maintained by third parties.

Running it yourself is the trade-off. Local rendering of a 33B video model on consumer hardware is currently slow — one clip can take a long stretch of GPU time, which breaks iterative workflows. Most creators pick one of two paths instead: a hosted API endpoint (fal.ai, per-second pricing below) or a browser generator like H3 Video, which runs MiniMax H3 in the cloud with nothing to install.

Resolution guide: 480p to 4K (and where H3 Max fits)

What resolution does MiniMax H3 support? Standard H3 generates up to 2K natively at 24fps, with 4K available on fal.ai. The H3 Maxvariant — post-trained by fal.ai for speed — tops out at 768p and hands 2K/4K and native-audio workloads back to standard H3. In practice the ladder looks like this:

TierAvailabilityTypical use
480pStandard H3 and H3 MaxFast drafts and volume production at the lowest per-second cost
768pStandard H3 and H3 Max (H3 Max max)Social clips and iteration; H3 Max renders 5s at 768p in under 3 seconds
2K nativeStandard H3Deliverable-grade output at 24fps — the tier H3 is known for
4KStandard H3 on fal.aiMaster-grade renders for post-production pipelines

A common workflow: prototype at 480p/768p on H3 Max for speed, then re-render the winning shot at 2K on standard H3. You get fast iteration and a high-resolution master without paying 2K prices on every attempt.

MiniMax H3 API cost

How much does the MiniMax H3 API cost? On fal.ai, MiniMax H3 is priced per second of generated video. Image-to-video on H3 Max is the cheapest entry point at $0.0125/s — a 15-second animated clip lands around $0.19 — while standard H3 at 2K and 4K commands the premium tiers:

Task480p768p2K4K
MiniMax H3 — text-to-video$0.05 / sec$0.06 / sec$0.13 / sec$0.16 / sec
MiniMax H3 Max — text-to-video$0.05 / sec$0.08 / sec——
MiniMax H3 Max — image-to-video$0.0125 / sec$0.02 / sec——

fal.ai also offers a free tier of up to 5 videos per day on the H3 Max model page. On H3 Video, generation is credit-based instead: you buy credit packs (no forced subscription), and the playground shows the exact credit cost before you generate — so the per-clip price is never a surprise.

The speed difference matters as much as the price: H3 Max renders a 5-second 768p clip in under 3 seconds — roughly 35x faster than the official MiniMax H3 endpoint — while standard H3 prioritizes fidelity over latency. For a deeper comparison, see the MiniMax H3 Max guide.

Capabilities you can use today

Four input modes

Cinematic camera control

H3 reads cinematography vocabulary directly: pan, zoom, tilt, and roll, plus timecode shot timing like [0s-3s] for a push-in followed by [3s-8s] for a slow orbit — so you can storyboard camera moves inside a single 15-second clip. For prompt technique, see the MiniMax H3 prompt guide with 25+ copy-paste examples.

Generate MiniMax H3 videos on H3 Video

H3 Video is an independent generator powered by the MiniMax H3 model: pick an input mode, write your prompt or upload media, see the credit cost, and generate — nothing to install, no local GPU, native audio on every clip. Start in the playground, or jump straight to the MiniMax H3 Max and H3 Max Turbo fast generators.

Create your first H3 AI video

Generate up to 15 seconds of 2K video with native stereo audio, powered by the MiniMax H3 model. See the credit cost before you generate — free to start.

Open the H3 Video generator

MiniMax H3 FAQ

Is MiniMax H3 open source?

MiniMax H3 ships with open weights: the 33B-parameter model is published on Hugging Face, so anyone can download it, run it locally, and fine-tune it. Local rendering on consumer hardware is currently slow, which is why most creators use a hosted endpoint or a browser generator like H3 Video instead.

Where can I download MiniMax H3 — Hugging Face or GitHub?

The official weights live on Hugging Face under the MiniMax organization — search "MiniMax H3" there. GitHub is where the community collects local-inference helpers, quantized builds, and fine-tuning tooling around the model.

What resolution does MiniMax H3 support?

Standard MiniMax H3 generates up to 2K natively at 24fps, with 4K available on fal.ai. The H3 Max variant is tuned for speed and tops out at 768p; 480p and 768p are the budget tiers for standard H3 as well.

How much does the MiniMax H3 API cost?

On fal.ai, standard MiniMax H3 text-to-video costs $0.05/s at 480p, $0.06/s at 768p, $0.13/s at 2K, and $0.16/s at 4K. H3 Max ranges from $0.0125/s (image-to-video, 480p) to $0.08/s (text-to-video, 768p). On H3 Video you buy credit packs and see the exact cost before every generation.

Does MiniMax H3 generate audio?

Yes — native stereo synchronized audio covering speech, sound effects, and ambience is generated together with the video, across the full clip duration.

How long can a MiniMax H3 video be?

Up to 15 seconds at up to 2K resolution and 24fps, with native audio across the full duration.

What is the difference between MiniMax H3 and MiniMax H3 Max?

H3 Max is a variant post-trained by fal.ai for speed: 5-second 768p clips in under 3 seconds, at the cost of the 2K/4K tiers and the focus on native audio. Use H3 Max to iterate fast and standard H3 for the final master.

Is H3 Video the official MiniMax H3 site?

No. H3 Video is an independent platform powered by the MiniMax H3 model. H3 Video is not affiliated with, endorsed by, or sponsored by MiniMax or Hailuo AI.

This page is an independent technical reference written by the H3 Video team and is not affiliated with, sponsored by, or endorsed by MiniMax, Hailuo AI, or fal.ai. MiniMax, Hailuo, and fal are trademarks of their respective owners. Specs, rankings, and per-second pricing reflect vendor-published information (MiniMax's H3 announcement and fal.ai's model pages) as of September 2026 and may change without notice.