Published · 2d agoProduct3 min read
DeepSeek's vision agent wins three of eleven benchmarks, all against May's Opus
V4-Flash-Vision-Exp is live on DeepSeek's API and costs a fraction of Anthropic's price. The vendor's own table shows it trailing on eight of eleven tests, including a 12-point gap on repository work.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- DeepSeek released an experimental multimodal model on Friday and said its agent performance is close to Anthropic's Opus-4.8.
- On DeepSeek's own published table, the new model beats Opus-4.8 on three benchmarks out of eleven.
- DeepSeek-V4-Flash-Vision-Exp is live on the company's API platform. It takes the text-only V4-Flash and adds the ability to read images and screenshots, and can then act on what it sees.
- The Hangzhou company also shipped version 0.1.1 of its agent harness, with support built in.
- DeepSeek wrote on X: "This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge".
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
DeepSeek put an experimental multimodal model on its API platform on Friday and said its agent performance is close to Anthropic's Opus-4.8 [1][3]. On the company's own published table it beats that model on three benchmarks out of eleven, and the comparison target is the Opus that arrived in May, not the Claude Opus 5 that Anthropic introduced on 24 July [2][7].
The product is real and shippable. DeepSeek-V4-Flash-Vision-Exp takes the text-only V4-Flash, adds the ability to read images and screenshots, and can act on what it sees; DeepSeek also shipped version 0.1.1 of its agent harness with support built in [3][4]. Bloomberg reported the release on Friday, describing Opus 4.8 as an advanced Anthropic model [10].
The arithmetic is where operators should spend their attention. Three wins out of eleven is a 27 percent hit rate, and the wins are narrow: DeepSWE by 1.3 points, Agents' Last Exam by 1.6, ZeroBench by 1.0 [11][22]. The eight losses include two wide ones. NL2Repo puts DeepSeek at 57.7 against 69.7, a 12-point gap, and DSBench-Hard at 63.6 against 71.7 [12]. Repository-scale work is exactly what enterprises are buying agents to do, so that is the number that decides deployments, not the near-ties on Toolathlon-Verified (75.9 to 76.2), Chartography (64.3 to 65.0) or Terminal Bench 2.1 (83.9 to 85.0) [13].
Opus-4.8 is not a straw man. Anthropic's model deprecation page lists it as Active, defined as fully supported and recommended for use, with no retirement earlier than May 2027, and five Opus models hold that status at once [8]. But the table carries no Opus 5 column, nobody else has published that comparison, and the release does not claim one, so performance against Anthropic's newest Opus is unknown [9].
Two disclosures in DeepSeek's own material deserve credit and adjust the headline. The advertised multimodal leap over V4-Flash on ApexBench (36.5 against 26.2) and Agents' Last Exam (27.3 against 25.2) is partly an artefact: DeepSeek's footnote says the text-only model "ignores multimodal elements contained therein" in those two evaluations, meaning a model without vision was scored on tests containing images [15][16]. Separately, the company undersells its text result. It says the vision model "matches" V4-Flash on text, but across the seven text benchmarks the vision variant wins six, with Toolathlon-Verified up 5.6 points, DeepSWE up 4.9 and DSBench-Hard up 4.0; Cybergym is the exception, where adding sight cost 1.4 points on a security benchmark [5][17][18]. Every figure is DeepSeek's, produced with its own harness in minimal mode at temperature 1.0 and top_p 0.95 [19].
Price is what makes a two-point deficit an argument rather than a loss. Research cited in the report found V4-Flash the cheapest well-known model to run, at about 87 cents per million words against roughly $50 from Anthropic, a spread of about 57 times [20][23].
What to watch: whether anyone publishes a V4-Flash-Vision-Exp against Opus 5 comparison [9]; whether the NL2Repo gap narrows in a non-experimental release; and what the shipped official V4-Pro, which DeepSeek says has significantly enhanced agent capabilities, scores on the same eleven tests [21]. One benchmark is worth watching for the field rather than the contest: on AutomationBench all three models sit in the mid-twenties, at 25.7, 25.1 and 27.2 [14].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
DeepSeek released an experimental multimodal model on Friday and said its agent performance is close to Anthropic's Opus-4.8.
ReportedView cited source - [2]
On DeepSeek's own published table, the new model beats Opus-4.8 on three benchmarks out of eleven.
ReportedView cited source - [3]
DeepSeek-V4-Flash-Vision-Exp is live on the company's API platform. It takes the text-only V4-Flash and adds the ability to read images and screenshots, and can then act on what it sees.
ReportedView cited source - [4]
The Hangzhou company also shipped version 0.1.1 of its agent harness, with support built in.
ReportedView cited source - [5]
DeepSeek wrote on X: "This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge".
ReportedView cited source - [6]
DeepSeek said that on multimodal agent benchmarks the model "makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8".
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- thenextweb.comAna Maria Constantin2d agoDeepSeek launches an experimental multimodal model to rival Anthropic
- siliconangle.comMaria DeutscheryesterdayDeepSeek debuts multimodal language model competitive with Opus 4.8



