Build1 distinct publisher3 min readPublished
AINews reports that fal post-trained Minimax's H3 and served it 35x faster than the official endpoint. Twitch and YouTube removed the resulting live stream immediately, so fal stood up its own player instead.
The Engineer · Build desk

build
A renderer that terminates itself is how an unwatched stream reports failure1 distinct publisher
build
Ng's new DeepLearning.ai spec is a reskilling checklist, not a hiring req1 distinct publisher
invest
AI video model renders faster than real time, enabling a 24-hour channel at a promotional $2,160 a day1 distinct publisher
product
From $50 to $2.25M: Niu Lai prices human authorship, not AI reputational risk2 distinct publishers
Compiled by The EngineerSomething wrong?How this is made
Thirty-five times against the official endpoint is a ratio, and a ratio needs a denominator you can inspect [1]. The AINews writeup does not supply one: no resolution, no frame budget, no batch size, no GPU count, no latency target for the reference endpoint [2]. A vendor's default inference path is normally tuned for cost across mixed traffic, which makes it the softest baseline in the building. The post-training and the engine are both fal's, so anyone reproducing this from the public H3 release starts two components short [1].
The relevant number is the point where frame rate crosses into watchable, not the multiplier by itself. Consistency models had already compressed a 30 second generation to about one second, a 30x cut [3][8], and one second per generation buys you one frame per second of wall clock [3]. Take 24 fps as the floor for something a person will sit through, which is my assumption and not the source's, and you need another 24x on top [9]. Fal's claimed 35x clears that with roughly 1.46x to spare [10]. That spare capacity is what "faster-than-realtime" is describing [7]. The two multipliers do not compose: the 30x was consistency-model work on earlier systems, the 35x is engine work on H3 [1][3].
The stream itself is where the product broke down. Once generation is continuous, the artifact is a stream, and a stream needs a player. Fal employees put theirs on Twitch and YouTube, and both platforms removed it immediately [4][5]. The account was removed without a cited policy provision or a stated reason from either platform, and there was no appeal [11]. This limit differs from the throughput ceiling fal had just solved: an engine rewrite can raise a throughput ceiling because fal controls the engine, but fal does not control the platforms' account decisions, and no configuration flag on fal's side reaches across that boundary.
Fal's response was to run its own live video service [5]. That means absorbing what the platforms had been providing at no charge: playback, moderation, abuse reporting, and egress for a feed that never ends. Those costs scale with viewer-hours rather than with generation, so the inference win does not offset them. Anyone planning a continuous generative-media product should price the distribution stack at design time, because it is now the component with a third-party veto attached.
On quality, AINews is blunt: pure slop, a mishmash with no plot and low-quality RL-tuned imagery, and nobody will watch it [6]. I will take the reviewer at their word. The load-bearing claim here is the frame rate, and it holds whether or not the content is watchable [7]. What it does not survive is a platform treating infinite generated video as its own category, which is a policy decision that no inference engine gets a vote in.
Ranked by verification strength, evidence, and original report placement.
Twitch and YouTube kicked fal off the platform immediately, so fal made its own live video service in the mould of 'twitch plays pokemon'.
The source states only that Twitch and YouTube removed fal immediately; it cites no policy provision, gives no reason from either platform, and reports no appeal.
The AINews item covering the fal result is a short narrative account and supplies no resolution, frame budget, batch size, GPU count, or latency figure for the official endpoint it compares against.
Before this, generating images and video took time: even using consistency models to get a 30 second generation down to 1 second left you with at best 1 FPS video, which AINews describes as well below anything acceptable for consumer-grade human attention.
The continuous generation result was first noticed by Ethan Mollick, then productized by fal employees into an infinite Twitch stream.
AINews describes the resulting stream as pure slop, a fever dream mishmash of content with no plot and low quality RL tuned imagery, and says nobody will actually watch it.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One newsletter, and no numbers behind the number
The 35x rests on a single clause in a single AINews issue. No resolution, no frame budget, no batch size, no GPU count, no baseline latency — nothing that would let a reader or a competitor check it. MiniMax, whose endpoint is the loser in this comparison, is not heard from. What is solidly evidenced is narrower: that a stream existed, that platforms pulled it, and that the writer watched it.
Two deployments, no audience
Something is genuinely running: a tuned model on fal's engine, a stream built on it, and a replacement player after the takedown. What is missing is anyone on the other end. No viewer count, no customer, no paid usage, no cost per minute — and the reporting itself predicts nobody will watch. Shipped, therefore, but not yet adopted by anybody outside the company that shipped it.
'Infinite video singularity' for a stream the writer calls slop
The framing runs well ahead of what is on the table. A barrier is declared broken on the strength of an unspecified multiple, in the same piece that says the output is a plotless mishmash nobody will watch. Our own arithmetic is what keeps the gap from being wider: 35x does clear the roughly 24x that a 1 FPS baseline needs to reach a watchable frame rate, with about 1.46 times to spare, so the direction of the claim is credible even where its size is unverified.
The multiple comes from the company selling the inference
Fal's business is the engine the speedup is attributed to, and the endpoint it beat belongs to the lab whose weights it borrowed — that is a benchmark chosen and run by the party it flatters. The relaying newsletter has no stake in fal, and to its credit says out loud that the demo is slop, but it also reprints the figure without asking for conditions. Nobody in this story had a reason to try to make the number smaller.
Confident about the events, not the magnitudes
We would stand behind the narrative spine: Mollick noticed, fal productized, the platforms cut it off, fal built its own player. The quantities are another matter — a 35x with no test conditions, inside a three-day roundup assembled from 12 subreddits and 544 Twitter accounts, is a figure we can report but not underwrite. One more account, from MiniMax or from fal with numbers attached, would move this sharply.