Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Your video API takes duration: 10 and renders 8.708 seconds

Causal 3D autoencoders quantise clip length to latent-frame boundaries, so only a short arithmetic progression of durations is reachable. Snap the number before you show, price or store it.

The Engineer · Build desk

How we use AISend a correction

What happened

  • The author cut a sequence of generated clips to a music bed and found it three frames out at the first transition, nine at the second, and by the sixth segment nothing lined up with anything.
  • Every generation was requested with duration: 10; the request returned 200 and produced an MP4 whose container duration was 8.708 seconds.
  • There was no warning field, no note in the response body, and no mention of the behaviour on the docs page the author had read three times.
  • The author states this is a general property of latent video models rather than a bug in one provider.
  • A video diffusion model works on a compressed latent tensor rather than frames, and the compression is temporal as well as spatial: a causal 3D autoencoder folds a run of input frames into a single latent frame.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer cutting six generated clips to a music bed found the first transition three frames out, the second nine frames out, and by the sixth segment nothing lined up with anything [1]. Every request had asked for ten seconds; the API returned 200 and an MP4 whose container duration was 8.708 seconds, with no warning field, no note in the response body, and no mention on the docs page he had read three times [3][15].

The author's argument is that this is a general property of latent video models rather than a bug in one provider [25]. A video diffusion model does not operate on frames but on a latent tensor compressed in time as well as space, and a causal 3D autoencoder folds a run of input frames into a single latent frame [16]. Because the encoder is causal, the first frame is kept whole and everything after it is compressed in groups, so with temporal stride s a clip of F frames becomes latent_frames = (F - 1) / s + 1 [17]. That divides evenly only when F is congruent to 1 modulo s; frame counts that miss the condition get padded or truncated, and implementations pick the nearest legal count and render that instead [18]. Stack the second constraint, that many of these models generate in fixed blocks of latent frames rather than one at a time, and the renderable lengths collapse to F = head + block * n [19]. Legal durations are those F values divided by the frame rate, and nothing between them is reachable [20]. A duration parameter, on this account, is a hint snapped to a grid the caller was never shown [21].

Three consequences, and only the first is visible. The UI promised ten seconds and the file is 8.708, so the UI lied [4]. The shortfall is 1.292 seconds per clip [24], which across six segments is roughly 7.75 seconds of drift [12]; the author puts it at about eight seconds, the difference between cutting on the beat and re-rendering the sequence [2]. Then billing: these APIs charge per second of output, so estimating from the requested duration while the model renders a longer legal block means quoting one number and charging another, which users find on their own [26][27].

The fix is to resolve duration to a legal value before anything is displayed, priced or persisted [5]. The published helper enumerates head + block * n clamped to the provider's min and max [22], then rounds requested seconds times fps to a target frame count and returns the nearest legal one along with its duration in seconds [23]. Three rules follow: show the resolved value in the duration control at the moment the user picks it, because a slider that snaps is honest and one that rounds in private is not [6]; price from frames / fps * rate, never from requested seconds [7]; and store the resolved value on the job, so the answer to "why do my six clips not add up" is a column rather than a reconstruction [8].

The integration test is the part worth copying. It sweeps requested durations from 4 to 9 seconds and ffprobes each result [9], and asserts not that the file matches the request, which the author says is a test you cannot pass and should not want, but that it matches the resolved duration the API reported [10][11]. Where a provider returns no resolved value, the author treats that absence as the finding, though the sentence is truncated in the published text [13].

Watch whether providers publish the grid parameters at all. The account names no provider and gives no frame rate [14], which means anyone integrating today has to map head, block and fps empirically with exactly that sweep, and re-check it whenever a model version changes.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption
Insufficient
Hype gap+12
Incentives55
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The author cut a sequence of generated clips to a music bed and found it three frames out at the first transition, nine at the second, and by the sixth segment nothing lined up with anything.

  2. [2]

    Second problem: the error compounds. Six segments each 1.3 seconds short is eight seconds of drift, which is the difference between cutting on the beat and re-rendering the sequence.

  3. [3]

    Every generation was requested with duration: 10; the request returned 200 and produced an MP4 whose container duration was 8.708 seconds.

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 21, 2026

    The duration your video API accepts is not the duration it renders

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories