We timed 21 production text to video endpoints on the same prompt, from the same machine, on the same afternoon. One of them renders five seconds of video in half a second.
| Metric | Value | What it measures |
|---|---|---|
| Model time | 0.49s | GPU render for 5 seconds of 480p video with synchronized audio |
| Generation time | 2.1s | Same request measured end to end on the platform, queue included |
| Peak throughput | 3.7x | Seconds of finished video per second of generation, 15s clip at 480p |
| Field median | 53.8s | Across the other 20 configurations tested the same afternoon |
Short answer: H3 Max Turbo is the fastest AI video generation model you can call from a public API today. It renders 5 seconds of 480p video with synchronized audio in 0.49 seconds of GPU time and returns the finished file in about 2 seconds. At 768p the same clip renders in 1.56 seconds and arrives in 3.8. No other endpoint among the 21 tested returned a finished clip in under 11 seconds on any run.
Speed claims for video models get argued past each other because three different numbers are all called "generation time". This test measured all three, and every table below says which one it is using.
Quote model time when you are comparing model architectures, generation time when you are comparing APIs, and round trip when you are deciding whether a person will sit and wait. The model time figures here were reproduced in a second, separate set of runs against the same endpoints, matching to within a few hundredths of a second at every setting.
H3 Max Turbo is the fastest AI video generation model available through a public API as of September 2026. It is a post-trained variant of the open weights MiniMax H3 model, built by fal and served on an inference stack designed around it. In this benchmark it generated a 5 second clip at 768p (1344 x 768, 24 fps, with sound) in a median 3.8 seconds, against 13.2 seconds for the next fastest endpoint tested.
| Model | Clip delivered | Median | Fastest run | vs real time | Audio | Published rate |
|---|---|---|---|---|---|---|
| H3 Max Turbo | 5.2 s at 1344x768 | 3.8 s | 3.3 s | 1.38x | Yes | $0.04 / s |
| H3 Max | 5.2 s at 1344x768 | 4.5 s | 4.3 s | 1.16x | Yes | $0.08 / s |
| LTX 2.5 Fast | 6.1 s at 1280x720 | 24.3 s | 22.4 s | 0.25x | Off in test | $0.09 / s |
| Veo 3.1 Lite | 4.0 s at 1280x720 | 28.8 s | 28.7 s | 0.14x | Off in test | $0.03 / s |
| Kling V3 Turbo Standard | 5.0 s at 1280x720 | 33.8 s | 33.5 s | 0.15x | Yes | $0.112 / s |
| Grok Imagine Video 1.5 | 5.0 s at 848x480 | 39.6 s | 33.7 s | 0.13x | Yes | $0.08 / s |
| Seedance 2.0 Mini | 5.0 s at 864x496 | 79.5 s | 67.0 s | 0.06x | Off in test | ~$0.072 / s |
| Hailuo 02 Standard | 5.9 s at 1366x768 | 83.2 s | 80.8 s | 0.07x | No | $0.045 / s |
| Wan 3.0 | 5.0 s at 1280x720 | 111.4 s | 110.8 s | 0.04x | Off in test | $0.10 / s |
| Wan 2.2 5B FastWan | 5.0 s at 1280x704 | 145.6 s | 23.2 s | 0.03x | No | $0.025 / clip |
Generation time for a finished clip at the 720p to 768p tier. Measured 7 September 2026, 3 samples per endpoint, all called through fal. Bars are scaled to the slowest median in this table.
The real time column divides the duration of the delivered clip by its generation time. A value above 1.0 means the model produced the video faster than the video plays. Rates are published list prices for the tier tested and exclude promotional discounts.
The same test at the entry resolution tier
Most video APIs offer a cheaper, smaller tier for drafts and iteration. That is where latency matters most, because it is where people press generate over and over. The ordering held.
| Model | Clip delivered | Median | Fastest run | vs real time | Audio | Published rate |
|---|---|---|---|---|---|---|
| H3 Max Turbo (480P) | 5.2 s at 832x480 | 2.1 s | 1.8 s | 2.47x | Yes | $0.025 / s |
| H3 Max (480P) | 5.2 s at 832x480 | 2.3 s | 2.1 s | 2.26x | Yes | $0.05 / s |
| PixVerse V6 (360p) | 5.0 s at 640x360 | 25.6 s | 23.2 s | 0.20x | No | $0.025 / s |
| Wan 2.2 5B FastWan (480p) | 5.0 s at 832x480 | 33.8 s | 11.3 s | 0.15x | No | $0.0125 / clip |
| Vidu Q3 Turbo (540p) | 5.0 s at 960x528 | 47.5 s | 35.0 s | 0.11x | Off in test | $0.035 / s |
| Kandinsky 5 Distill | 5.0 s at 768x512 | 51.4 s | 45.9 s | 0.10x | No | $0.05 / clip |
| LongCat Video Distilled (480p) | 5.0 s at 832x480 | 56.3 s | 56.1 s | 0.09x | No | $0.005 / s |
| Wan 2.2 A14B Turbo (480p) | 5.0 s at 832x480 | 87.7 s | 18.8 s | 0.06x | No | $0.05 / clip |
Generation time at each model's entry resolution, 3 samples per endpoint.
Two things stand out in the entry tier. First, H3 Max and H3 Max Turbo land within a fifth of a second of each other at 480p, because at that size the work around the render dominates the render itself. The gap opens at 768p, where Turbo's model time was 1.56 s against 2.56 s for standard H3 Max. Second, several endpoints show a wide spread between their fastest and their median run, which is the signature of capacity and cold starts rather than of model architecture.
Distilled and research endpoints
Distilled checkpoints are the usual answer to "what is the cheapest way to make a video", so they were tested too, with acceleration set to the most aggressive option each one exposes.
| Model | Clip delivered | Median | Fastest run | Audio |
|---|---|---|---|---|
| LTX-2 19B Distilled | 4.9 s at 1024x576 | 13.2 s | 11.9 s | Yes |
| LTX Video 13B 0.9.8 Distilled | 5.0 s at 1280x704 | 15.2 s | 12.7 s | No |
| Cosmos Predict 2.5 Distilled | 5.8 s at 1280x704 | 80.1 s | 74.3 s | No |
| Infinity Star | 5.1 s at 1280x720 | 96.5 s | 44.8 s | No |
| LTX 2.3 22B Distilled | 5.0 s at 1024x576 | 144.4 s | 134.5 s | Yes |
| LongCat Video Distilled (720p) | 10.9 s at 1280x704 | 274.7 s | 262.7 s | No |
Distilled and research text to video endpoints, generation time, 2 samples each. Wide spreads indicate cold starts on lower traffic endpoints.
A video model is faster than real time when the time to deliver N seconds of finished video is less than N seconds. Ask for a 10 second clip, get it back in 7 seconds, and you have crossed the line. It is the threshold that separates a render job from an interactive experience, because below it a person is always waiting on the machine, and above it the machine is waiting on the person.
H3 Max Turbo clears that line across most of its configuration space, on all three clocks up to 1080p.
| Setting | Delivered | Model time | Generation | Round trip | vs real time |
|---|---|---|---|---|---|
| 480P, 15 s | 15.1 s at 832x480 | 1.72 s | 4.0 s | 4.7 s | 3.74x |
| 480P, 5 s | 5.2 s at 832x480 | 0.49 s | 2.1 s | 3.1 s | 2.47x |
| 768P, 10 s | 10.1 s at 1344x768 | 4.36 s | 6.5 s | 7.1 s | 1.55x |
| 768P, 5 s | 5.2 s at 1344x768 | 1.56 s | 3.8 s | 4.3 s | 1.38x |
| 768P, 15 s | 15.1 s at 1344x768 | 8.56 s | 12.3 s | 13.1 s | 1.23x |
| 1080P, 5 s | 5.2 s at 1920x1080 | 2.24 s | 4.8 s | 5.7 s | 1.08x |
| 2K, 5 s | 5.2 s at 2560x1440 | 4.23 s | 8.5 s | 9.4 s | 0.61x |
H3 Max Turbo text to video, all three clocks by setting. One sample per row, measured 7 September 2026. Ratios are against generation time.
Every row above came back with synchronized audio. The 1080P and 2K settings are latent refinements of a native 768p generation, and even 2560x1440 was delivered in under 10 seconds. Published per second rates cover the 480p and 768p native modes, so check the model page for the current rate at the refined resolutions.
Image to video, on the same three clocks
The image to video endpoint behaves the same way, measured here from a 16:9 input image.
| Setting | Delivered | Model time | Generation | Round trip | vs real time |
|---|---|---|---|---|---|
| 480P, 15 s | 15.1 s at 832x480 | 1.85 s | 5.1 s | 5.8 s | 2.93x |
| 480P, 5 s | 5.2 s at 832x480 | 0.49 s | 3.1 s | 4.0 s | 1.67x |
| 768P, 15 s | 15.1 s at 1344x768 | 8.63 s | 13.4 s | 14.4 s | 1.13x |
| 768P, 5 s | 5.2 s at 1344x768 | 1.70 s | 5.2 s | 6.1 s | 1.00x |
H3 Max Turbo image to video from a 1280x720 input, 2 samples per row, measured 7 September 2026.
One practical finding worth knowing before you optimise anything else: image to video output follows the aspect ratio of the input image, and that moves both the clock and the bill. The identical 768p 15 second request from a square 1024x1024 input rendered at 768x768, roughly half the pixels of 16:9, and its model time fell from 8.63 s to 3.47 s. The shape of the still you feed in is a latency lever.
Three choices compound, and none of them is simply "a smaller model".
The Turbo variant pushes that further with a shorter render path, which is why its model time at 768p came in roughly 40 percent below the standard H3 Max endpoint on an identical request.
Among the models in this benchmark, H3 Max Turbo has the lowest published per second rate that includes synchronized audio: $0.04 per second of video at 768p and $0.025 per second at 480p. That is $0.20 for a five second 768p clip, $0.60 for a fifteen second one, and $2.40 per finished minute.
Note: A promotional launch rate of $0.01 per second at 768p and $0.00625 per second at 480p runs until 14 September 2026.
| Model | Per second | 5 second clip | Per minute | Native audio |
|---|---|---|---|---|
| H3 Max Turbo | $0.04 | $0.20 | $2.40 | Yes |
| Veo 3.1 Lite | $0.05 | $0.25 | $3.00 | Yes |
| H3 Max | $0.08 | $0.40 | $4.80 | Yes |
| LTX 2.5 Fast | $0.09 | $0.45 | $5.40 | Yes |
| Wan 3.0 | $0.10 | $0.50 | $6.00 | Yes |
| Kling V3 Turbo Standard | $0.112 | $0.56 | $6.72 | Yes |
| Grok Imagine Video 1.5 | $0.14 | $0.70 | $8.40 | Yes |
| Seedance 2.0 Mini | ~$0.155 | ~$0.77 | ~$9.28 | Yes |
Published list price per second of generated video at the 720p to 768p tier, audio included where the model supports it.
A fair caveat, because "cheapest" depends on what you are counting. Several distilled research endpoints publish lower per second rates: LongCat Video Distilled lists at $0.01 per second at 720p, and LTX Video 13B Distilled at $0.02 per second. They return silent video, and in this test their median generation times were 274.7 seconds and 15.2 seconds. If your metric is dollars per second of finished, watchable, audio bearing video from a model that also sits at the top of the public preference boards, H3 Max Turbo is the value leader in this field. If your metric is the absolute floor on a per second rate for silent draft footage, the distilled endpoints deserve a look.
One cost effect appears on no price list: you are not billed for waiting, but your team is. A creative iterating through ten variations saves roughly eight minutes of dead time per round at 2 seconds a clip instead of 54.
Speed invites the question of what was traded away, so it is worth reading the independent boards.
Design Arena summarised the combination as delivering "the quality of MiniMax H3 at more than 50x the speed". Batuhan Taskaya, Head of Engineering at fal, framed the goal this way at launch: "Generative video has historically forced developers to choose between quality and speed. By pairing the new capabilities of H3 Max and fal's post training we've proven that this tradeoff is no longer necessary."
H3 Max Turbo is a text to video and image to video model with the following envelope, read from its live API schema on 7 September 2026:
How to call the API
// npm i @fal-ai/client
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("minimax/h3-max-turbo/text-to-video", {
input: {
prompt: "A golden retriever puppy runs across a sunlit kitchen floor toward the camera, warm morning light through the window, shallow depth of field",
resolution: "480P", // 480P | 768P | 1080P | 2K
duration: 5, // 5 to 15 seconds
aspect_ratio: "16:9",
prompt_expansion_mode: "balanced"
}
});
console.log(result.data.video.url);
console.log(result.data.timings.inference); // model time, in seconds
# pip install fal-client
import fal_client
result = fal_client.subscribe(
"minimax/h3-max-turbo/text-to-video",
arguments={
"prompt": "A golden retriever puppy runs across a sunlit kitchen floor toward the camera",
"resolution": "480P",
"duration": 5,
"aspect_ratio": "16:9",
"prompt_expansion_mode": "balanced",
},
)
print(result["video"]["url"])
The endpoint is serverless and pay per use, with no minimum spend and no GPU to provision. Signed in users also get five free generations a day of up to 15 seconds each with audio through the fal sandbox, which is the quickest way to check the numbers in this article for yourself.
Latency is one axis among several, and the right pick depends on what the clip is for.
Reproducibility matters more than the headline number, so here is exactly what was measured.
Third party endpoints fluctuate with demand, so treat any single latency figure as a snapshot. The ordering at the top of the table was stable across every run.
What is the fastest AI video generation model?
H3 Max Turbo is the fastest AI video generation model available through a public API as of September 2026. It renders 5 seconds of 480p video with synchronized audio in 0.49 seconds of GPU time and returns the finished file in about 2 seconds. No other endpoint among the 21 tested returned a finished clip in under 11 seconds on any run.
How fast is H3 Max Turbo exactly?
Three clocks answer that, all measured on 7 September 2026. Model time, the GPU render reported per request: 0.49 seconds for a 5 second 480p clip and 1.56 seconds at 768p. Generation time, measured by the platform across the whole request: 2.1 seconds and 3.8 seconds. Round trip from a laptop, including the network hop: 3.1 seconds and 4.3 seconds.
Why do published speed numbers for the same model disagree?
Because three different things get called generation time. The lowest is model time, the render pass on the GPU, which is 1.56 seconds for a 5 second 768p clip. Generation time adds queue admission, prompt expansion, encoding and audio muxing, at 3.8 seconds. Round trip adds the network hop from your own machine, at 4.3 seconds. All three are real, and the response body returns model time in a timings.inference field so you can separate them yourself.
Is AI video generation real time yet?
For clips up to 1080p, yes. H3 Max Turbo generated a 15 second 480p clip in 4.0 seconds, which is 3.7 seconds of finished video per second of waiting, and even a 5 second 1080p clip came back in 4.8 seconds. Fully continuous, steerable video is a separate product category, served by session based models such as H3 Max Director rather than by request and response endpoints.
How fast is H3 Max Turbo image to video?
From a 16:9 input image, generation times were 3.1 seconds for a 5 second 480p clip, 5.1 seconds for a 15 second 480p clip, 5.2 seconds at 768p for 5 seconds and 13.4 seconds at 768p for 15 seconds. Output follows the aspect ratio of the input image, so a square still renders roughly half the pixels of 16:9 and returns proportionally faster.
What is the cheapest AI video model per second?
Among the models benchmarked here, H3 Max Turbo has the lowest published list rate that includes synchronized audio, at $0.04 per second of video at 768p and $0.025 per second at 480p, which is $2.40 per finished minute at 768p. Distilled research endpoints such as LongCat Video Distilled publish lower rates near $0.01 per second, but they return silent video and took far longer to deliver in this test.
Does H3 Max Turbo generate audio?
Yes. Every generation returns a synchronized AAC audio track produced in the same pass as the picture, covering room tone, foley, ambience and music. Describing the sound in the same prompt as the shot is what steers it, and there is no separate audio step afterwards.
What resolutions and durations does H3 Max Turbo support?
Durations run from 5 to 15 seconds in a single request. Resolutions are 480P and 768P natively, plus 1080P and 2K produced as latent refinements from a 768p source. At 16:9, 768P is 1344 x 768 at 24 fps, and text to video supports 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16.
What is the difference between H3 Max Turbo and H3 Max?
Both are post-trained variants of MiniMax H3 with the same resolution and duration envelope. Turbo takes a shorter render path, which made it about 15 percent faster on generation time at 768p in this test and roughly 40 percent faster on model time, and it lists at half the per second price. Standard H3 Max is the choice when you want the longer render on the same architecture.
Can I generate AI video for free?
Yes. Signed in fal users get five free generations a day of up to 15 seconds each with synchronized audio through the sandbox, on a rolling 24 hour window with no subscription. Past that allowance the same endpoints are pay per use through the API.
Which AI video model is best for speed and quality together?
H3 Max leads the public preference boards while running faster than real time: Design Arena scores it first on image to video at an Elo of 1,341, and Artificial Analysis ranks it first on image to video with audio at an Elo of 1,201 over 2,177 samples. That combination is what makes the family the practical pick when both latency and output quality matter.
How was this benchmark measured?
Each endpoint was called with an identical prompt from one machine on 7 September 2026, recording the platform's own generation time for the request, the model time where the endpoint reports it, and the wall clock round trip. Every output file was probed with ffprobe to confirm its resolution, frame rate, duration and audio stream, and medians are reported alongside the fastest observed run from H3 Max model page.