Build1 distinct publisher3 min readPublished
Ant Ling's documentation samples video at two frames a second, caps clips at 30 seconds and then keeps only 32 frames, leaving the frame budget as the real constraint a GUI agent has to plan around, well short of the million-token context window.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Multiply the documented video limits out and the frame ceiling arrives first. Thirty seconds sampled at two frames per second is 60 frames, and the documentation keeps 32 of them, which puts full-rate coverage at 16 seconds [9][1]. So either you send a 16-second clip and keep every sampled frame, or you send the full 30 and accept that something upstream throws away half of them. The frame budget, not the million-token context window described in the architecture [8], is the binding limit on video work.
The image path has its own envelope. Requests take JPEG and PNG, up to 40 images, inside a 32 MB maximum request body [10]. Fill all 40 slots and each image gets 0.8 MB on average [2]. An agent that carries a scrollback of screenshots in one request is therefore trading image count against image sharpness, and small interface text is the first thing to go when you downscale. Ant Ling's own framing points the same way: the company's claims about searching a timeline and extracting keyframes appear to rely on an agent splitting the job into tool calls rather than putting a long recording into one request [15].
The 38 on the Intelligence Index belongs to the base model, and Artificial Analysis measured it, along with the 5.1 billion active parameters per token [3][4]. The 42 belongs to Ant Ling, which attributes the gain to training visual and language capabilities together [5]. As of September 4, Artificial Analysis had published no page for the VL model [6]. For those four points to transfer to your own shortlist, the party that produced the 38 would have to score the VL model on the same index at the same version, and the mechanism would have to be that image training moved a text score. That is a testable claim, and on this evidence nobody outside Ant has tested it.
The sparsity is the part of the design that looks well aimed. About 4 percent of the parameters do work on any given token [3], which is a serving-cost story for whoever holds the weights and a latency story for everyone else. Whether you can be the former is not settled by this release: it documents the vision transformer, the two-layer MLP projector and VideoRoPE positional encoding [7], the hybrid backbone [8] and the request limits [9][10], and it states no licence or weights download for Ling-3.0-flash-VL [4]. Ant Group's open-source record under Richard Bian [16] provides useful context, though it stops short of a licence file for this release.
The demonstration that will get quoted is the one where the model takes a design reference, writes the site, opens it in a browser, compares the render against the reference and revises the code [11]. Ant Ling also says a single reference image can steer layout, color, typography and interface components across a new site [13]. These are company demonstrations, not independent evaluations [14]. The loop they describe spends one screenshot per iteration, which places the 40-image, 32 MB envelope directly across the feature the release is built around, and that is the first number I would test rather than the index score.
Ranked by verification strength, evidence, and original report placement.
On September 4, Ant Ling released Ling-3.0-flash-VL, a version of its fast language model that accepts images and video and can act on what it sees.
Ling-3.0-flash-VL is built on Ling-3.0-flash, the 124 billion-parameter mixture-of-experts model Ant introduced in July 2026.
Artificial Analysis independently gives Ling-3.0-flash a score of 38 on its Intelligence Index.
Artificial Analysis measures 5.1 billion active parameters per token for Ling-3.0-flash.
Ant Ling says the visual version raises the Intelligence Index score to 42, attributing the improvement to training visual and language capabilities together.
Artificial Analysis had not published a corresponding page for the VL model as of September 4, so the four-point gain remains Ant Ling's claim.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
GPT-5.6 Luna scores 52 against a peer median of 17. Its token count is what lands on your bill1 distinct publisher
build
Spark 1.3's index jump lands on the three tests that carry half the score6 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor documentation plus one outside score
Nearly every specific traces to Ant Ling: the architecture, the demo reel, the request caps. The exception is Artificial Analysis's 38 and 5.1 billion active parameters, and those describe the text base, not the visual model this release is about. Runtimewire labels the demos as company demonstrations and marks the unverified score, which keeps this in mid-range rather than lower, though every test run so far has stayed inside Ant.
Availability only, one day in
The whole record is availability: an OpenAI-compatible endpoint, published integration instructions, and a two-week hosted trial priced from $20. Users, customers and third-party integrations have yet to appear, and the frame-budget analysis reads the documentation rather than anyone's production traffic. With no weights on offer, independent testers cannot yet add to this.
A million tokens, thirty-two frames
The million-token window is the figure in the launch copy; 32 frames is the figure in the developer documentation, and for video work the second one governs. Ant's four-point Intelligence Index gain sits in the same position: asserted at launch, unmeasured on the visual model. Runtimewire surfaces both gaps instead of repeating the framing, so the overstatement belongs to Ant's presentation rather than to this reporting.
Launch narrated by its seller
The story is assembled from a launch by the party that profits from it, with Ant Ling's product and growth lead as the human thread. Access is paid and hosted, which conveniently keeps the inference environment the browser and GUI demos depend on inside Ant's stack, and the licence question that would let outsiders retest the claimed score stays open. Artificial Analysis, the one party with nothing to sell here, has scored only the older text model.
Firm on limits, thin on quality
Documented ceilings and the stated architecture are stable ground, and the arithmetic on frames, sparsity and per-image size follows straight from them. What the release actually buys a developer turns on whether joint visual-language training moved the model four points, and that has one source with a stake in the answer. Runtimewire's reading of the timeline and keyframe claims is plausible but its own inference.