fal released H3 Max, a post-trained version of MiniMax’s H3 video model, on August 26, 2026, and the number that matters is a ratio: a 5-second clip in under 3 seconds of wall time. The model renders video faster than the video plays, and by two independent leaderboards it did not give up quality to get there. MiniMax lists it on its own API as MiniMax-H3-Max at $0.08 a second for 768p.
This page does two things. First, the record: what shipped, what it costs, what the leaderboards say, and what the model still cannot do, all from fal’s and MiniMax’s own pages. Second, the argument: why crossing the realtime line is the specific thing that lets an AI employee stop being text on a screen.
On this page · 7 sectionsOpen
- H3 Max is a post-trained version of MiniMax’s open-weight H3 video model, developed by fal Research and released August 26, 2026, jointly listed by MiniMax as MiniMax-H3-Max. It is not a MiniMax-only release, and ‘H3 Turbo’ is Artificial Analysis’s internal name for it, not the product name.
- The headline claim, in fal’s words: ‘generating a 5-second video under 3 seconds of wall time,’ about 35x the throughput of the official MiniMax H3 endpoint and on average 15x faster than models of comparable quality. fal’s own API reports roughly 2.5 seconds of backend inference for a 5-second 768p clip.
- Quality did not pay for the speed, by the leaderboards: #1 on Artificial Analysis’s image-to-video board with audio (Elo 1,201, 2,177 samples) and #1 on Design Arena’s image-to-video board (Elo 1,341, ahead of the base H3 at 1,333). fal’s internal Bayesian Elo study against 12 models, including Seedance 2.5, Kling 3, and Veo 3.1, puts it first on all three dimensions.
- Specs: 480p or 768p (1344x768 at 24 fps), 5 to 15 second clips, text-to-video and image-to-video with first and last frames, native synchronized audio. No 2K, and MiniMax’s docs say no reference-to-video on the Max variant; both stay on the standard H3.
- Price: $0.05 per second at 480p and $0.08 per second at 768p on both MiniMax’s API and fal’s, so a 5-second 768p clip costs $0.40. fal’s launch discount ends September 7. fal’s own 15-second cost table: H3 Max $1.20, Wan 3.0 $1.50, Kling v3 $1.89, FLUX 3 $2.55, Seedance 2.5 $7.10.
- The unlock for AI employees is not cheaper marketing video. It is that a video reply now takes about as long as a text reply: a standing worker with a reference face and a cloned voice can answer with a talking clip inside a shift, at the speed of the conversation.
- CellCog’s video tools run on Seedance 2.5 today and its avatars already carry a face reference and a cloned voice. Faster-than-realtime models are what would turn that pairing into a face that talks back, and this page tracks the models that get there.
§ 01What shipped
MiniMax launched H3 on July 31, 2026 as a general-purpose multimodal generation model: video with native stereo audio, up to 15 seconds at 2K, with a promise to open the weights. fal took those weights, post-trained them with new data aimed at prompt adherence and visual quality, and built an inference engine around the result. fal’s announcement is careful about what that means: “Traditional step-distillation methods aim to match their base model while we aim for a better model at much faster speed.”
| Item | Detail |
|---|---|
| Released | August 26, 2026, by fal Research; MiniMax lists it as MiniMax-H3-Max |
| Base model | MiniMax H3 (open weights, July 31, 2026) |
| Headline speed | 5-second clip in under 3 seconds of wall time; ~2.5 s backend inference at 768p (fal) |
| Throughput claim | ~35x the official MiniMax H3 endpoint; ~15x models of comparable quality (fal) |
| Resolution | 480p or 768p; 768p is 1344x768 at 24 fps; no 2K |
| Duration | 5 to 15 seconds, whole seconds; a 15-second clip renders in about 15 seconds |
| Modes | Text-to-video (six aspect ratios), image-to-video with optional last frame |
| Not supported | Reference-to-video (image, video, or audio reference) and 2K, both on standard H3 only |
| Audio | Native, synchronized, generated in the same pass |
| Price, MiniMax and fal | $0.05/s at 480p, $0.08/s at 768p, output only |
| Launch discount | fal: 75% off, ends September 7, 2026 |
| Free tier | fal sandbox: five generations a day, up to 15 seconds each |
§ 02The speed, with the receipt
fal’s exact wording: “It does so while generating a 5-second video under 3 seconds of wall time, which is roughly 35x the throughput of the official MiniMax H3 endpoint and on average 15x faster than anything with comparable quality.” The product page adds the mechanism for checking it: every API response carries a timings.inference field, the actual denoising time on the backend, “which lands at roughly 2.5 seconds for a 5-second 768p generation.” fal’s published example response on the image-to-video endpoint shows 2.77 seconds.
That is the claim, and it is a testable one, which is more than most video launches offer. Two honest edges to it. The wall time a user sees includes upload, queueing, and download, so “under 3 seconds” describes the model, not your app’s round trip. And the ratio holds at 5 seconds; at 15 seconds fal says the clip takes about 15 seconds, so the model runs at roughly realtime on the longest setting and roughly twice realtime on the shortest.
§ 03The quality, by people who are not fal
fal ran its own head-to-head preference study against twelve models, including the official MiniMax H3 endpoint, Gemini Omni Flash, Wan 3.0, Seedance 2.5, Kling 3, and Veo 3.1, scoring overall preference, prompt understanding, and aesthetics with Bayesian Elo ratings and 95% confidence intervals. H3 Max ranks first on all three. That is a vendor study, and we treat it as one.
The independent boards say the same thing, which is why the speed number is worth a page.
| Board | H3 Max result | Reference point |
|---|---|---|
| Artificial Analysis, image-to-video with audio | #1, Elo 1,201 (95% CI plus or minus 11, 2,177 samples), listed as “MiniMax H3 Turbo (768p)” | Ahead of Seedance 2.0, MiniMax H3, Gemini Omni Flash, Grok Imagine Video 1.5, Veo 3.1, Kling 3.0 per fal’s release |
| Design Arena, image-to-video | #1, Elo 1,341 | Base MiniMax H3 at 1,333 |
The Design Arena gap over the base model is eight Elo points: a post-train that is faster and slightly better, rather than faster and worse. That is the part fal’s announcement calls moving the frontier, and on these two boards it is a fair description.
§ 04The price, and the cost table fal wants you to read
Standard rates match on both APIs. MiniMax’s pay-as-you-go page lists MiniMax-H3-Max at $0.05 a second for 480p and $0.08 for 768p, billed on output video only; standard H3 is $0.08 at 768p and $0.13 at 2K. fal’s launch discount, 75% off on its model pages as of this writing, ends September 7.
| Model and resolution | Price per second | 5-second clip | 15-second clip |
|---|---|---|---|
| MiniMax-H3-Max, 480p | $0.05 | $0.25 | $0.75 |
| MiniMax-H3-Max, 768p | $0.08 | $0.40 | $1.20 |
| MiniMax-H3, 768p | $0.08 | $0.40 | $1.20 |
| MiniMax-H3, 2K | $0.13 | $0.65 | $1.95 |
Same price as the standard model at 768p, about 35 times the throughput. The trade is 2K and reference-to-video, which stay on standard H3.
fal also publishes a cross-model cost table for 15 seconds of roughly 720p video at fal’s own August 26 rates, with Kling quoted on its cheapest v3 tier that generates audio, since H3 Max always does. It is a vendor’s table on a vendor’s platform, and prices elsewhere differ, but the ordering is the point.
| Model | 15 seconds |
|---|---|
| MiniMax H3 Max, 768p | $1.20 |
| Wan 3.0, 720p | $1.50 |
| Kling v3 Standard with audio, 720p | $1.89 |
| FLUX 3, 720p | $2.55 |
| Seedance 2.5, 720p | $7.10 |
§ 05The line that matters: generation faster than playback
Here is why a video model earns a page on a blog about AI employees.
Every AI employee today is text. It writes the report, files the task, sends the email, refreshes the dashboard. Video, when it appears at all, is a production step: a script, a render that takes minutes, a review, a publish. That shape means video is something an employee makes occasionally, the way a human employee occasionally records a screencast. It is not something an employee is.
Faster-than-realtime generation changes the shape. When a 5-second clip renders in under 3 seconds, a video reply costs about what a paragraph costs in wall time. An employee with a reference face and a cloned voice can answer a question with a short talking clip, in the thread, while the question is still warm. The owner’s morning status can be a face saying it. A customer’s how-do-I question can get a fifteen-second demonstration instead of a wall of steps. None of that is a marketing video. It is the ordinary work of a standing role, now with a face on it.
Three properties have to hold at once for that to be real, and H3 Max is the first model to publish credible numbers on all three:
- Speed past realtime. Under 3 seconds for 5 seconds of video, with the timing exposed in the API response.
- Quality good enough to represent someone. First on two independent image-to-video boards; a face that drifts or a mouth that fails to sync is worse than text.
- A price that survives frequency. $0.40 for a 5-second reply at 768p, so a worker that answers in video a dozen times a shift is spending single-digit dollars, not a production budget.
What is missing is identity control. MiniMax’s API reference is explicit that MiniMax-H3-Max does not support reference-to-video: you cannot hand it a reference image, video, or voice clip to lock a character. Image-to-video with a first frame is the workaround, and fal’s own examples show one character holding across six locations in a single generation, so the model can hold a face. But a standing employee needs its face locked across thousands of generations over months, not six shots in one, and that is the reference-to-video job. When it lands on the Max variant, or on a competitor at the same speed, the third property is met and the unlock is complete.
§ 06What this means for CellCog
CellCog’s video generation runs on Seedance 2.5 today, chosen because it accepts dozens of reference images, videos, and audio clips per segment and assembles multi-minute films from one prompt: the production job, done well. CellCog avatars already carry the other half of the pairing, a reference face and a cloned voice, and every avatar video ships disclosed as AI-made.
H3 Max is the first model whose numbers make the third thing plausible: a face that answers at conversational speed. It is not there yet, for the reason above, and neither is any competitor at this speed. So the honest position is a watch, not a switch. The signals this page updates on: reference-to-video arriving on H3 Max or on any faster-than-realtime model, independent replication of the under-3-seconds figure outside fal’s own timing field, and the September 7 end of fal’s discount. If a model clears the bar, CellCog employees get a face the same way they got Gemini 3.8 Flash and Fable 5.1: underneath the role, on the day it is ready, with nothing to change on your side. A full shift of real work runs about $25, and the cost depends purely on how much work you assign.
§ 07The record
As of September 3, 2026: H3 Max is available on fal (text-to-video and image-to-video endpoints, free sandbox at five a day) and on MiniMax’s pay-as-you-go API as MiniMax-H3-Max; list price $0.05/s at 480p and $0.08/s at 768p; fal’s launch discount ends September 7; no 2K and no reference-to-video on the Max variant; #1 on Artificial Analysis image-to-video with audio (1,201) and Design Arena image-to-video (1,341). If reference-to-video ships on H3 Max, if a competitor publishes a comparable faster-than-realtime figure with comparable rankings, or if the prices move, the update happens here, same day.
Q1Is H3 Max a MiniMax model or a fal model?
Both, honestly stated. MiniMax released H3 as an open model on July 31, 2026, and said it would open the weights. fal Research post-trained those weights with new data aimed at prompt adherence and aesthetics, co-designed its inference engine around the result, and released H3 Max on August 26. MiniMax lists it on its own API and pricing page as MiniMax-H3-Max, and fal’s press release describes it as jointly released. fal calls itself the original creator of the Max variant.
Q2Is the speed claim independently verified?
Partly. The under-3-seconds figure is fal’s, and fal exposes it: every API response carries a timings.inference field, and fal’s published example shows 2.77 seconds for a 5-second image-to-video generation. Anyone with an account can reproduce that. The quality rankings are independent: Artificial Analysis and Design Arena both place it first on image-to-video, and fal’s own 12-model preference study is the vendor’s.
Q3What can it not do?
No 2K output, no clips shorter than 5 seconds or longer than 15, and, per MiniMax’s API reference, no reference-to-video on the Max variant: you cannot yet pass a reference image, video, or audio clip to steer identity. Image-to-video with a first frame is the closest substitute, and fal’s examples show a single character holding across six locations inside one generation. Those limits are exactly the ones to watch; this page updates when they move.
Q4How does it compare to Seedance 2.5?
fal’s preference study ranks H3 Max ahead of Seedance 2.5 on overall quality, prompt understanding, and aesthetics, and fal’s cost table prices a 15-second 768p H3 Max clip at $1.20 against $7.10 for Seedance 2.5 at 720p on fal. Seedance 2.5 keeps advantages H3 Max does not claim: reference-to-video with up to dozens of reference files, longer segments, and the multi-shot film protocols built around it. Our Seedance 2.5 pricing guide has the full picture.
Q5Does CellCog use MiniMax H3 Max?
Not today. CellCog’s video generation runs on Seedance 2.5, chosen for reference-to-video and multi-minute film assembly, and CellCog avatars already carry a reference face and a cloned voice. Every faster-than-realtime model is evaluated against that stack, and if H3 Max or a successor earns a place, the switch happens underneath your employees without a change on your side. A full shift of real work runs about $25, and the cost depends purely on how much work you assign.
