Product1 distinct publisher3 min readUpdated
V4-Flash-Vision-Exp is live on DeepSeek's API and costs a fraction of Anthropic's price. The vendor's own table shows it trailing on eight of eleven tests, including a 12-point gap on repository work.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
DeepSeek put an experimental multimodal model on its API platform on Friday and said its agent performance is close to Anthropic's Opus-4.8 [1][3]. On the company's own published table it beats that model on three benchmarks out of eleven, and the comparison target is the Opus that arrived in May, not the Claude Opus 5 that Anthropic introduced on 24 July [2][7].
The product is real and shippable. DeepSeek-V4-Flash-Vision-Exp takes the text-only V4-Flash, adds the ability to read images and screenshots, and can act on what it sees; DeepSeek also shipped version 0.1.1 of its agent harness with support built in [3][4]. Bloomberg reported the release on Friday, describing Opus 4.8 as an advanced Anthropic model [10].
The arithmetic is where operators should spend their attention. Three wins out of eleven is a 27 percent hit rate, and the wins are narrow: DeepSWE by 1.3 points, Agents' Last Exam by 1.6, ZeroBench by 1.0 [11][22]. The eight losses include two wide ones. NL2Repo puts DeepSeek at 57.7 against 69.7, a 12-point gap, and DSBench-Hard at 63.6 against 71.7 [12]. Repository-scale work is exactly what enterprises are buying agents to do, so that is the number that decides deployments, not the near-ties on Toolathlon-Verified (75.9 to 76.2), Chartography (64.3 to 65.0) or Terminal Bench 2.1 (83.9 to 85.0) [13].
Opus-4.8 is not a straw man. Anthropic's model deprecation page lists it as Active, defined as fully supported and recommended for use, with no retirement earlier than May 2027, and five Opus models hold that status at once [8]. But the table carries no Opus 5 column, nobody else has published that comparison, and the release does not claim one, so performance against Anthropic's newest Opus is unknown [9].
Two disclosures in DeepSeek's own material deserve credit and adjust the headline. The advertised multimodal leap over V4-Flash on ApexBench (36.5 against 26.2) and Agents' Last Exam (27.3 against 25.2) is partly an artefact: DeepSeek's footnote says the text-only model "ignores multimodal elements contained therein" in those two evaluations, meaning a model without vision was scored on tests containing images [15][16]. Separately, the company undersells its text result. It says the vision model "matches" V4-Flash on text, but across the seven text benchmarks the vision variant wins six, with Toolathlon-Verified up 5.6 points, DeepSWE up 4.9 and DSBench-Hard up 4.0; Cybergym is the exception, where adding sight cost 1.4 points on a security benchmark [5][17][18]. Every figure is DeepSeek's, produced with its own harness in minimal mode at temperature 1.0 and top_p 0.95 [19].
Price is what makes a two-point deficit an argument rather than a loss. Research cited in the report found V4-Flash the cheapest well-known model to run, at about 87 cents per million words against roughly $50 from Anthropic, a spread of about 57 times [20][23].
What to watch: whether anyone publishes a V4-Flash-Vision-Exp against Opus 5 comparison [9]; whether the NL2Repo gap narrows in a non-experimental release; and what the shipped official V4-Pro, which DeepSeek says has significantly enhanced agent capabilities, scores on the same eleven tests [21]. One benchmark is worth watching for the field rather than the contest: on AutomationBench all three models sit in the mid-twenties, at 25.7, 25.1 and 27.2 [14].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
DeepSeek released an experimental multimodal model on Friday and said its agent performance is close to Anthropic's Opus-4.8.
On DeepSeek's own published table, the new model beats Opus-4.8 on three benchmarks out of eleven.
DeepSeek-V4-Flash-Vision-Exp is live on the company's API platform. It takes the text-only V4-Flash and adds the ability to read images and screenshots, and can then act on what it sees.
The Hangzhou company also shipped version 0.1.1 of its agent harness, with support built in.
DeepSeek wrote on X: "This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge".
DeepSeek said that on multimodal agent benchmarks the model "makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8".
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but wholly vendor-sourced
There is an unusual amount of numeric detail — eleven itemised benchmark results, disclosed harness mode and sampling settings, and a vendor footnote about the blind text baseline — and the model's API availability is a checkable fact. But every capability figure originates with DeepSeek, run on DeepSeek's own harness, and no independent or third-party replication exists in the supplied material. The single-publisher cluster also means no second outlet has verified the table.
Shipped and callable, no usage evidence
Concrete availability signals exist: the experimental model is live on DeepSeek's API, agent harness 0.1.1 ships with support, and the official V4-Pro has shipped. But the supplied source reports no deployments, customer names, download or traffic figures, or any usage disclosure, and the vision model is explicitly labelled experimental. Adoption is therefore availability only, with the cost spread cited as a reason buyers are looking rather than evidence that they have switched.
Framing outruns the table
DeepSeek's public framing — a major leap bringing multimodal agent performance close to Opus-4.8 — sits above what its own table shows: three wins in eleven, eight losses, and a 12-point deficit on NL2Repo, the repository-scale work enterprises buy agents for. The headline multimodal jump is partly an artefact of scoring a text-only baseline on tests containing images it cannot see. The baseline is also May-vintage, with no Opus 5 comparison available. The gap is moderate rather than severe because DeepSeek disclosed the footnote itself, said 'close to' rather than 'ahead of', never claimed to beat Anthropic's newest model, and actually understates its own text-benchmark gains.
Vendor-run scoreboard, vendor-chosen baseline
DeepSeek authored the benchmarks, ran them in its own harness, selected which competitor model appeared as a column, and chose which of its own models to omit — the officially shipped V4-Pro is absent from the table. The price comparison that gives the release its commercial force also favours DeepSeek. Mitigating factors are the disclosed sampling settings and the self-supplied footnote on the blind baseline. The reporting outlet also cites its own prior coverage twice, a mild self-reference incentive.
Single publisher relaying one vendor's figures
Fact-level confidence in the release itself is high — the model is live, the harness version shipped, Bloomberg covered the launch the same day, and Anthropic's lifecycle page is a checkable third-party artefact. Confidence in the capability comparison is much lower: one publisher, one vendor's self-run numbers, an experimental model, no independent replication, and the decisive Opus 5 comparison unrun. The overall reading is therefore held down by concentration of sourcing rather than by internal inconsistency.
build
A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026