Build1 distinct publisher3 min readPublished
The five-second clip in three seconds is fal's own figure, measured on fal's hardware against an endpoint it does not control, and the throughput multiple only transfers if you know the batch size and GPU count behind it.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Two numbers, two axes. The latency figure is clip length over wall clock: five seconds of output in about three seconds, roughly 1.7x faster than playback [3][14]. The throughput figure is about 35 times MiniMax's official H3 endpoint [4]. Those do not measure the same property. If the 35x were a per-request latency comparison on identical hardware, MiniMax's endpoint would need something near 105 seconds for the same five-second clip [15]. The more ordinary explanation is batch depth, GPU count and queue policy, and fal's announcement does not say which [4][5]. Throughput without a denominator is a claim about someone else's capacity plan.
For the multiple to reach your bill, several things have to hold: same GPU class, same resolution and frame count, same sampling step count, the audio track in the same state, and enough concurrent requests to fill whatever batch fal is filling [11]. Traffic that arrives one request at a time from a UI collects the three seconds and none of the 35x.
The gate is the better engineering here. fal's inference group changed precision, sampling and execution, and kept a change only when fal's evaluations showed output quality held up [7]. Every knob in that list buys speed with fidelity, so the gate is load-bearing rather than ceremonial. The evaluations were head-to-head human preference studies covering overall quality, prompt understanding and aesthetics [6]. That is self-grading, and unavoidably so: once fal Research adds training data to the base checkpoint [2], there is no public model with the same identity left to score against.
The commercial mechanism follows from the same fact. When two hosts expose the same weights, competition collapses onto latency, throughput, uptime and price [8]. Change the weights and design the runtime around them, and a broader GPU platform cannot reproduce the endpoint by loading the public checkpoint [9]. That is the wall fal is building against Replicate, Modal, RunPod and CoreWeave [10], and it is a real wall rather than a positioning statement. It also marks a move up the stack from hosting other labs' models to shaping them [17].
It is paid for twice. fal pays in per-model engineering: post-training plus a runtime pass, per base model, which does not amortise across a catalogue the way generic serving does [2][7]. The customer pays in portability. H3 Max exists as fal's text-to-video and image-to-video endpoints [11], and fal's platform also distributes MiniMax's own models [19], so the exit from H3 Max leads back to weights that were deliberately changed [2], with prompts tuned against the modified ones.
Scope matters on the ranking too. The first-place claim covers specific image-to-video categories, not the broader text-to-video leaderboard [12], and the source material cuts off before naming the benchmark in full [13]. The comparison against 12 leading video models was internal [5]. Two endpoints also means two integrations, and the placement only speaks to one of them [11][12].
Adoption therefore turns on a measurement fal cannot run for you: p95 latency at your own concurrency, on your own prompts, with audio in the configuration you ship [11]. The three-second figure is cheap to test in an afternoon. The 35x needs fal's hardware and MiniMax's rate limits, so it stays a claim about their workload until someone publishes the denominator.
Ranked by verification strength, evidence, and original report placement.
fal, based in San Francisco and co-founded by Burkay Gur and Gorkem Yurtseven, released H3 Max on September 1, post-trained from the open-weights MiniMax H3 video model.
fal Research added training data intended to improve prompt adherence and visual quality, while fal's inference engineers optimized the serving stack alongside the evolving model.
In a September 1 release, fal said H3 Max can generate a five-second video in roughly three seconds.
According to fal's technical announcement, H3 Max delivers about 35 times the throughput of MiniMax's official H3 endpoint.
fal compared H3 Max with 12 leading video models in internal testing, and the results remain a fal-reported comparison rather than a standardized independent speed benchmark.
fal's researchers evaluated post-training checkpoints through head-to-head human preference studies covering overall quality, prompt understanding and aesthetics.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Fal reached continuous video generation by tuning inference on someone else's model1 distinct publisher
invest
AI video model renders faster than real time, enabling a 24-hour channel at a promotional $2,160 a day1 distinct publisher
build
Harness choice moved token use 83-fold with the model held constant1 distinct publisher
build
H3's reference path is a different checkpoint, capped at 12 files, and stops at 768p1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One announcement, relayed once
Three seconds, 35x, first place: all three trace to fal's own September 1 announcement, and Runtimewire names PR Newswire as where it came in. The single figure fal did not choose — third at 1,235 Elo on the Artificial Analysis text-to-video board — reaches us secondhand and is flagged as a dated snapshot. Nothing in this story has been run on hardware fal does not control.
Endpoints live, traffic unstated
What is real is that you can call it: two endpoints, in production, the day the claim was made. Beyond that the scale numbers belong to a different subject and a different year — the two million developers and 300-plus enterprise customers come from fal's July 2025 Series C and describe the platform, not this model. No one has said how much H3 Max traffic exists.
Superlatives cropped to fit
The pitch is a 35x multiple against the endpoint of the lab that wrote the weights, travelling without batch size or GPU count — the two things that decide whether it means anything to a single request. 'First on independent benchmarks' survives only inside image-to-video; on text-to-video the model sits third. The overstatement lives in fal's framing rather than in the coverage, because Runtimewire does most of the puncturing itself.
Vendor sells it, vendor scores it
fal is selling the endpoint, wrote the evaluation, picked the twelve models it beat, chose the prompts and the human judges, and pushed the result out over a paid wire. The strategic argument in the piece — that identical public weights leave only speed and price to compete on — is precisely the argument that pays for a large multiple, whatever the multiple measures.
Solid on what was said, thin on what is true
There is little doubt about what fal announced or how it says it got there; the announcement is detailed and the reporting labels its own limits. Whether H3 Max is 35 times faster than anything remains open. An independent latency run at a stated batch size, or a current Artificial Analysis entry, would move this in either direction quickly.