Build1 publisherNot yet confirmed elsewhere3 min readPublished
Your video API takes duration: 10 and renders 8.708 seconds
Causal 3D autoencoders quantise clip length to latent-frame boundaries, so only a short arithmetic progression of durations is reachable. Snap the number before you show, price or store it.
The Engineer · Build desk
What happened
- The author cut a sequence of generated clips to a music bed and found it three frames out at the first transition, nine at the second, and by the sixth segment nothing lined up with anything.
- Every generation was requested with duration: 10; the request returned 200 and produced an MP4 whose container duration was 8.708 seconds.
- There was no warning field, no note in the response body, and no mention of the behaviour on the docs page the author had read three times.
- The author states this is a general property of latent video models rather than a bug in one provider.
- A video diffusion model works on a compressed latent tensor rather than frames, and the compression is temporal as well as spatial: a causal 3D autoencoder folds a run of input frames into a single latent frame.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer cutting six generated clips to a music bed found the first transition three frames out, the second nine frames out, and by the sixth segment nothing lined up with anything [1]. Every request had asked for ten seconds; the API returned 200 and an MP4 whose container duration was 8.708 seconds, with no warning field, no note in the response body, and no mention on the docs page he had read three times [3][15].
The author's argument is that this is a general property of latent video models rather than a bug in one provider [25]. A video diffusion model does not operate on frames but on a latent tensor compressed in time as well as space, and a causal 3D autoencoder folds a run of input frames into a single latent frame [16]. Because the encoder is causal, the first frame is kept whole and everything after it is compressed in groups, so with temporal stride s a clip of F frames becomes latent_frames = (F - 1) / s + 1 [17]. That divides evenly only when F is congruent to 1 modulo s; frame counts that miss the condition get padded or truncated, and implementations pick the nearest legal count and render that instead [18]. Stack the second constraint, that many of these models generate in fixed blocks of latent frames rather than one at a time, and the renderable lengths collapse to F = head + block * n [19]. Legal durations are those F values divided by the frame rate, and nothing between them is reachable [20]. A duration parameter, on this account, is a hint snapped to a grid the caller was never shown [21].
Three consequences, and only the first is visible. The UI promised ten seconds and the file is 8.708, so the UI lied [4]. The shortfall is 1.292 seconds per clip [24], which across six segments is roughly 7.75 seconds of drift [12]; the author puts it at about eight seconds, the difference between cutting on the beat and re-rendering the sequence [2]. Then billing: these APIs charge per second of output, so estimating from the requested duration while the model renders a longer legal block means quoting one number and charging another, which users find on their own [26][27].
The fix is to resolve duration to a legal value before anything is displayed, priced or persisted [5]. The published helper enumerates head + block * n clamped to the provider's min and max [22], then rounds requested seconds times fps to a target frame count and returns the nearest legal one along with its duration in seconds [23]. Three rules follow: show the resolved value in the duration control at the moment the user picks it, because a slider that snaps is honest and one that rounds in private is not [6]; price from frames / fps * rate, never from requested seconds [7]; and store the resolved value on the job, so the answer to "why do my six clips not add up" is a column rather than a reconstruction [8].
The integration test is the part worth copying. It sweeps requested durations from 4 to 9 seconds and ffprobes each result [9], and asserts not that the file matches the request, which the author says is a test you cannot pass and should not want, but that it matches the resolved duration the API reported [10][11]. Where a provider returns no resolved value, the author treats that absence as the finding, though the sentence is truncated in the published text [13].
Watch whether providers publish the grid parameters at all. The account names no provider and gives no frame rate [14], which means anyone integrating today has to map head, block and fps empirically with exactly that sweep, and re-check it whenever a model version changes.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The author cut a sequence of generated clips to a music bed and found it three frames out at the first transition, nine at the second, and by the sixth segment nothing lined up with anything.
- [2]
Second problem: the error compounds. Six segments each 1.3 seconds short is eight seconds of drift, which is the difference between cutting on the beat and re-rendering the sequence.
- [3]
Every generation was requested with duration: 10; the request returned 200 and produced an MP4 whose container duration was 8.708 seconds.
- [4]
First problem: the output is not the length promised. The UI said 10s and the file is 8.708s, so the UI lied, though not in a way anyone notices on one clip.
- [5]
The author's fix is to stop treating duration as a free variable at the edge of the system and to resolve it to a legal value before anything is displayed, priced or persisted.
- [6]
Rule one: show the resolved duration instead of the requested one, changing the duration control at the moment the user picks it, because a slider that snaps is honest and a slider that accepts anything and rounds in private is not.
- [7]
Rule two: price as frames / fps * rate from the resolved frames, never from the requested seconds.
- [8]
Rule three: store the resolved value on the job, so that when someone asks why six clips do not add up the answer is in a column rather than a reconstruction.
- [9]
The author's integration test loops over requested durations 4, 5, 6, 7, 8, 9 and 10, creates a job with the prompt "a still grey card", downloads the file and ffprobes it.
- [10]
The test asserts that the probed file duration is close to job.resolvedDuration to two decimal places, not that it matches the request.
- [11]
The author says a test that the file matches the request is one you cannot pass and should not want; matching the API's reported resolved value is a contract you can hold a provider to.
ReportedSupportedSource: dev.to post by the author2 sources— create a free account to open themView cited source - [12]
Six clips each 1.292 seconds short accumulate about 7.75 seconds of drift.
- [13]
The author writes that if the provider returns no resolved value at all, the absence is itself the finding; the published sentence is cut off mid-word.
- [14]
The account does not name the provider and does not state the frame rate or the grid's head, block, min and max values.
- [15]
There was no warning field, no note in the response body, and no mention of the behaviour on the docs page the author had read three times.
- [16]
A video diffusion model works on a compressed latent tensor rather than frames, and the compression is temporal as well as spatial: a causal 3D autoencoder folds a run of input frames into a single latent frame.
- [17]
Because the encoder is causal, the first frame is kept whole and everything after it is compressed in groups; with a temporal stride of s, a clip of F frames becomes latent_frames = (F - 1) / s + 1.
- [18]
The latent-frame formula divides evenly only when F is congruent to 1 modulo s; frame counts that miss that condition get padded or truncated, so implementations pick the nearest legal count and render that instead.
- [19]
Many of these models generate in fixed blocks of latent frames rather than one at a time, so the set of renderable lengths collapses into the arithmetic progression F = head + block * n for natural n.
- [20]
Every legal duration is one of the legal F values divided by the frame rate, and nothing between them is reachable.
- [21]
The author characterises duration: 10 as not a request but a hint that gets snapped to a grid the caller was never shown.
- [22]
The published legalFrames function enumerates frame counts head + block * n, breaking above the provider's max and including only those at or above its min.
- [23]
The published quantise function rounds requested seconds times fps to a target frame count, reduces the legal frame list to the nearest value, and returns the requested seconds, the resolved frames and frames / fps as seconds.
- [24]
The per-clip shortfall between the requested and rendered duration is 1.292 seconds.
- [25]
The author states this is a general property of latent video models rather than a bug in one provider.
ReportedInsufficientSource: dev.to post by the author2 sources— create a free account to open themView cited source - [26]
These APIs bill per second of output.
- [27]
Third problem: estimating cost from the requested duration while the model renders a longer legal block means quoting one number and charging another, which users discover themselves.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toThe duration your video API accepts is not the duration it renders
1 article · August 21, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Integration TestingFollow
- API contract designFollow
- Latent Video Diffusion InternalsFollow
- Usage-based billingFollow
- Video Generation APIsFollow
Entities
- MiniMax H3Follow
- MiniMaxFollow
- minimax-h3ai.videoFollow
- ffprobeFollow
- Causal 3D AutoencoderFollow
- TypeScriptFollow