Build1 distinct publisher3 min readPublished
DeepSeek V4 Flash Vision Exp undercuts Gemini 3.7 Flash by 3.4x on input tokens and 5.7x on output, then spent 3,467 completion tokens and 30.5 seconds on an invoice both models audited correctly.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
DeepSeek's tokenizer counted a test image at roughly 500 prompt tokens; Google's counted the same image at roughly 1,150 [13], and neither number shows up on either company's price page. Run those through the input rates and one page costs about $0.00011 on DeepSeek against about $0.00086 on Gemini, near eight times cheaper before the model has said a word [2]. For extraction work where the answer is four fields and a flag, that ratio is the entire business case.
The output side runs the other way. On the invoice audit DeepSeek billed 3,467 completion tokens for an answer of a few paragraphs, which the reviewer at The New Stack reads as internal reasoning billed as output [12][14]. That is a lot of deliberation about a monitor arm. Price the call with image tokens standing in for input: DeepSeek lands near $0.0024, Gemini near $0.0044 [3]. So a sticker gap of 3.4x on input and 5.7x on output [1] compresses to about 1.8x on that call, and at the doubled weekday rate DeepSeek's version costs roughly 9 percent more than Gemini's [3].
The sticker ratio only reaches your invoice if two conditions hold. Completion length has to stay bounded, which means capping output tokens and accepting that a reasoning-heavy model will sometimes hit the cap mid-answer. And the traffic has to sit outside the peak window. Both prices here are gateway prices on OpenRouter, which DeepSeek's model only reached on August 27 [2][5].
Clock time tracks the tokens: 30.5 seconds against 7.9 on the same audit [12], about 3.9x [4]. The reviewer calls that the only anomaly across the three tests [12], so the other calls were closer. Google markets Gemini 3.7 Flash as its "most intelligent workhorse model yet" [7]; on the two tasks whose results are reported, the workhorse tied a budget model on accuracy and won on latency.
The accuracy read is thinner than the pricing read. On the dual-axis chart, where revenue ran 0 to 12 on the left and costs 0 to 16 on the right so the cost line never visually clears the bars, both models answered Q1, named Subscriptions, and estimated $36.1 million, matching the source data [9]. On the invoice DeepSeek went further and noticed that the printed subtotal did not reconcile with the printed line items [11]. That is one graded call per model per test, by one grader, with the images and prompts published for replication [17][16]. That is more replication material than most comparisons ship; it is still too few runs to support a latency budget.
The test set is also cleaner than production. A synthetic invoice with three planted arithmetic errors and a chart with a deliberately mismatched second axis are legible inputs [8][9]. In my experience vision models diverge on skewed phone photos and stamps printed over totals, and nothing here measures that.
Where that leaves the buy: if the pipeline is a queue that tolerates a 30-second tail and runs off-peak, the discount is real and the input tokenization makes it larger than advertised. If a human is waiting on the screen, or the batch runs inside the peak window, the cheaper model is not reliably the cheaper call.
Ranked by verification strength, evidence, and original report placement.
The model keeps the budget V4 Flash price of $0.22 per million input tokens and $0.66 per million output tokens.
DeepSeek's price for the model doubles during weekday peak hours.
The reviewer sent the same images and prompts to both models through OpenRouter across three tests imitating back-office work, recording accuracy, tokens, cost and speed for each call.
On the invoice test DeepSeek took 30.5 seconds and billed 3,467 completion tokens, while Gemini answered in 7.9 seconds with 944 completion tokens; the reviewer describes this as the only anomaly of the three tests.
DeepSeek counted each test image at roughly 500 prompt tokens and Gemini at roughly 1,150, which the reviewer takes as evidence the two companies tokenize images very differently.
The reviewer attributes DeepSeek's 3,467 completion tokens on a short answer to a large amount of internal reasoning billed as output.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
DeepSeek V4 doubled its OpenRouter token share, and the bill it displaced was ~130x bigger1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
build
Harness choice moved token use 83-fold with the model held constant1 distinct publisher
invest
DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One tester, one run, receipts attached
The measurements are unusually legible for a vendor comparison — per-call tokens, seconds and dollars, with the prompts and images published so anyone can repeat them. What holds the score down is arithmetic, not honesty: three hand-built images, one call each, no second run to tell a 30.5-second outlier from noise, and a single publisher behind every figure. The cost conclusions also rest on approximate image token counts and on estimating input from images alone.
On the gateway, nobody's stack yet
What we can actually observe is availability: two releases eight days apart and a DeepSeek endpoint reachable through OpenRouter by August 27. Beyond that the story is silent — no request volumes, no customer names, no production rollouts, not even a claim that anyone is running invoices through this at scale. The only usage documented anywhere here is a reviewer spending about a cent and a half.
Cheap until Monday morning
The gap sits in the framing rather than the reporting. A 3.4x and 5.7x list-price advantage sounds decisive; by the time you account for reasoning tokens billed as output and the weekday doubling, the same invoice call lands about 9 percent above Gemini. Meanwhile both vendors' workhorse claims survive the test only because nothing separated the models on accuracy — every planted trap failed against both. The reviewer flags the weekend caveat himself, which keeps this a modest overstatement instead of a large one.
Two pitches quoted, one weekend bill paid
The pressure here is mostly the vendors': a document-understanding pitch and a superlative about workhorse intelligence, both reproduced without contest, both serving companies with obvious reasons to be compared favourably on cost. DeepSeek's own pricing structure is the sharpest interested party in the story — a headline rate that doubles when business hours arrive. The reviewer's stake looks small and is partly disclosed, including the admission that weekend testing is DeepSeek's best case.
Re-runnable, not yet re-run
Confidence rides on reproducibility rather than corroboration. The prompts and images are on the table, the token and cost figures reconcile against the published three-test totals, and the arithmetic is checkable. But no second publisher has touched any of it, one run per call leaves the latency spread unexplained, and our own earlier reading of the log test was wrong about what the reporting contained — a small caution about how much weight a single excerpt should bear.