Invest1 publisher3 min readPublished
DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores
The V4.1-Flash change log puts Terminal Bench 2.1 at 90.6 against V4-Pro's 87.9 and DeepSWE at 74.2 against 62.7. DeepSeek says API prices came down with the release and points to a pricing page for the amounts.
The Investor · Invest desk

What happened
- DeepSeek released V4.1-Flash on 10 September 2026, calling it the smallest model in a new architecture family and giving it native multimodal visual understanding on the API.
- The published scores include GPQA Diamond at 90.9, a Codeforces rating of 3471 and MathArena Apex at 65.6, all measured by DeepSeek itself.
- V4 Flash and V4 Flash Vision Exp are retired, with both of their model names temporarily routed to V4.1 Flash for compatibility.
- DeepSeek said it will keep serving V4 Pro on the API after 14 September 2026 with billing unchanged, citing user demand.
- API prices were reduced alongside the release, and the change log sends users to a separate Models and Pricing page for the amounts.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint Anyone who pinned deepseek-v4-flash has lost the ability to pin a version: the name still resolves, but to a different model, and DeepSeek describes the routing as temporary.
- decision A buyer benchmarking Chinese inference against a US vendor should now do it at the Flash tier, because that is where the agent scores of the previous flagship have landed.
- cost The cut arrives in the release note without a number, so per-token spend cannot be modelled from the announcement; the pricing table has to be pulled first.
- precedent An experimental vision branch was absorbed into the mainline model within three weeks of shipping. That is the lifespan to expect from anything DeepSeek labels experimental.
DeepSeek's V4-Pro general release of 13 August listed Terminal Bench 2.1 at 87.9, NL2Repo at 61.5, CyberGym at 83.3, Agents' Last Exam at 25.7 and HLE with tools at 60.0 [7]. The V4.1-Flash entry of 10 September lists the same five at 90.6, 65.4, 88.1, 31.8 and 63.9 [5][6]. Eight measures appear in both entries, and the smaller model leads on seven [3]. DeepSWE moves the furthest, 62.7 to 74.2, though the newer entry labels it v1.1 [5].
The exception is the closed-book exam. On HLE without tools, Pro scores 42.7 and Flash 36.8, a gap of 5.9 points, or 3.6 against the 39.1 DeepSeek reports for the pure-text subset alone [4][7][6]. Flash leads where tools do the retrieval and trails where the model has to hold the answer itself [3][6].
Three versions of one benchmark appear in the new entry: the same model scores 90.6 on Terminal-Bench 2.1, 30.0 on 3.0 and 31.2 on 4.0 [5]. The 60.6-point drop between 2.1 and 3.0 measures how much harder the newer test is [7]. So the 2.7-point margin over Pro on 2.1 sits at the top of a scale where both models score near 90 [4].
DeepSeek said prices came down with the release and pointed to its Models and Pricing page for the amounts [10]. The numbers live on that page. The release is consistent with cheaper Chinese inference without measuring it; the one price ratio in the recent entries dates from 13 August, when DeepSeek set off-peak rates at half of peak [11].
Every score here is DeepSeek's own. The August footnote gives the conditions: DeepSeek Harness minimal mode, max effort level, topp 0.95, temperature 1.0 [14]. That same entry said the vision model's leap on visual agent benchmarks brought "its multimodal agent capabilities close to Opus-4.8", which is DeepSeek's comparison and not an independent one [13].
The retirements say more about where DeepSeek is spending. V4-Flash-Vision-Exp shipped on 21 August as an experimental model and was pulled 20 days later, its name routed temporarily to V4.1-Flash [12][8][2]. Vision now sits in the mainline Flash model [1]. V4 Pro, meanwhile, keeps its API service past 14 September with billing unchanged because users asked for it [9].
I would run the price comparison at the Flash tier. A buyer weighing per-token cost against a US vendor is now comparing against a small model that carries the agent scores the Pro tier posted 28 days earlier [1][3]. If the pricing page shows only a token cut, this is an SKU consolidation with a benchmark table attached [10]. And if the seven wins are specific to DeepSeek's own harness, independent runs will put Flash back below Pro on tool-using agent tasks [14].
What to watch
- The Models and Pricing page: the per-token cut alongside V4.1-Flash is what turns this benchmark table into a cost comparison.
- Independent evaluations of V4.1-Flash on Terminal-Bench 2.1 and DeepSWE, run outside DeepSeek's own harness settings.
- Whether the temporary routing of the deepseek-v4-flash names gets a hard cutoff date, and whether V4 Pro's reprieve holds past it.