Skip to content

Build2 publishers2 min readPublished

Cohere says Embed 5 Fast keeps 98.4% of Pro's retrieval score when querying a Pro index

Cohere's Embed 5 lets teams index with Pro at $0.12 per million text tokens and query with Fast at $0.08 against the same vectors. The quality and throughput figures behind that split come from Cohere's own tests, run on datasets and parsing pipelines a buyer may not share.

The Engineer · Build desk

Illustration accompanying Cohere says Embed 5 Fast keeps 98.4% of Pro's retrieval score when querying a Pro index

What happened

  • Cohere reports that Fast processed documents at an average of 2.4 times Pro's throughput in its own tests.
  • Cohere introduced RCP-nDCG@10, a scoring method in which an AI judge grades documents against query-specific relevance criteria instead of fixed labels.
  • Embed 5 shipped on September 30 through the Cohere API, Model Vault, Microsoft Foundry and Amazon SageMaker, with single-tenant deployment via Model Vault.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Since the query encoder can change without re-embedding the corpus, a team can launch on Pro end to end and move the query path to Fast later, or back, as an ordinary deploy.
  • cost At Cohere's recommended 1,024-dimension int8 setting, 100 million chunks need about 102 GB, one eighth of the float32 footprint, and Fast can still query those vectors.
  • constraint The rival comparisons ran on text parsed by ViDoRe's authors or by Gemini 1.5 Flash, so a team with a different parsing pipeline cannot assume the ranking holds on its own records.

Pro writes the document vectors. Fast encodes each query into the same space at the same dimensions, and the vector store compares the two directly [1][6].

On Cohere's numbers, that swap costs 1.6 points against a Pro-to-Pro baseline of 100 [3]. Running Fast on both sides costs 3.4 points. Indexing with Pro therefore buys back 1.8 of them, a little over half [3]. Cohere says no individual dataset showed a major drop under the mix [5]. It scored the cross-model test with standard nDCG@10, not with its new metric [15].

Fast's text price is a third below Pro's, or $80 per billion text tokens against $120 [1][2]. Image inputs cost $0.40 per million tokens on either model, so an image query path gets no discount [2]. A $40 gap per billion tokens will rarely be the line that gets this change through review [2].

Throughput is the stronger case. Cohere's 2.4x figure is document throughput from its own tests [7]. Cohere assigns Fast to interactive search, high-volume retrieval and agent workflows [8]. Those workloads wait on the latency of one short query. Bulk document throughput is a separate measurement. The announcement does not include a customer workload or an independent production result [18].

The rival margins are narrow. On ViDoRe V3, Pro leads Voyage 4 Large by 2.1 points and Fast leads it by 0.8 [5]. On Cohere's parsed-PDF suite, Pro scored 84.8 to 83.6 for Voyage 4 Large and 80.8 for Gemini Embedding 2 [12]. The multilingual results are mixed. Pro leads a five-language European average with 77 to Gemini Embedding 2's 73, but trails Gemini on nine of the individual tests [16].

Cohere says it validated the RCP-nDCG@10 judge against human judgment, and it also used the method to optimize the models being scored [14]. The metric grades reranking over a fixed candidate set. First-stage retrieval from the full corpus is scored separately, with standard nDCG and Recall [15]. The New Stack notes that the reported scores are not directly comparable across Cohere's evaluations [15].

For any of this to transfer, a team's queries would need to look like the 40 datasets in the cross-model test, and its documents would need to reach the embedder in a similar state [3]. The shared space makes that cheap to check. One Pro index, both encoders on the query side and the team's own labelled queries are enough to measure the 1.6-point figure on the team's own data [1][3].

What to watch

  • An independent evaluation of Embed 5 against Voyage 4 Large and Gemini Embedding 2 on standard nDCG, using page images instead of parsed text.
  • A production report giving per-query latency for Fast against a Pro-built index, which Cohere's document-throughput figure does not measure.
  • Whether evaluators outside Cohere adopt RCP-nDCG@10, or its AI-judge scores stay confined to Cohere's own results.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories