Cohere launches Embed 5 with a cheaper model for live enterprise search

The Pro and Fast models share an embedding space, letting customers index with one and run queries with the other without rebuilding their index.

By · Published

Primary source: X

Why it matters

Cohere is making enterprise retrieval a cost-and-latency choice: customers can index with Pro and query with the cheaper Fast model in the same embedding space. Its performance claims rely on company-run evaluations, including a new Cohere scoring method.

Paper fragments flow into two different arrangements of points, representing Cohere Embed 5 Pro and Embed 5 Fast.

Cohere released Embed 5 on September 30th, a pair of embedding models aimed at enterprise search and retrieval systems: Pro for retrieval quality and Fast for lower-latency, higher-volume queries. The models share an embedding space, so customers can build an index with Pro and query it with Fast. Cohere co-founder and CEO Aidan Gomez (@aidangomez), one of the co-authors of the 2017 paper "Attention Is All You Need," is leading a company still competing on the practical infrastructure around enterprise AI, not only on generative models. Cohere announced Embed 5 in a thread on X.

Photo attached to Cohere's Embed 5 announcement post on X.
Photo attached to Cohere's Embed 5 announcement on X. Original image from @cohere.

The shared space is the product's central operating proposition. Embedding models turn text, images, or other inputs into vectors that software can search by meaning. Companies use them to find relevant records before a generative model drafts an answer or an AI agent takes an action. Embed 5 Pro costs $0.12 per million text tokens; Fast costs $0.08. Image inputs cost $0.40 per million tokens for either model, according to Cohere's product announcement.

The pairing gives customers a potential way to reserve the more expensive model for indexing large collections and use the cheaper one on the repeated query path. Cohere says the two versions can be mixed without rebuilding the index, and recommends indexing with Pro and querying with Fast. That is a potentially useful cost and deployment option; the published price alone does not establish total savings, which depend on workload, data volume, latency needs and downstream search infrastructure.

Cohere reports that Pro scored 85.8 on ViDoRe V3, compared with 83.7 for Voyage 4 Large and 83.2 for Google's Gemini Embedding 2. Fast scored 84.5. These are company-reported results, not an independent test. Cohere also introduced RCP-nDCG@10, a retrieval-scoring method used in its evaluations. The method grades documents against query-specific relevance criteria with an AI judge, rather than relying only on pre-existing relevance labels. Cohere says it validated the method against human judgment, but its own explanation describes the evaluation as a new approach the company developed and used to optimize its models. Comparisons using it deserve that context.

The benchmark details narrow the claim further. Cohere says the ViDoRe V3 documents were parsed as text and that ViDoRe's authors curated the evaluation material. For a separate parsed-document suite, Cohere reports an average of 84.8 for Pro, ahead of Voyage 4 Large at 83.6 and Gemini Embedding 2 at 80.8; the documents in that suite were parsed using Gemini 1.5 Flash. Those results support a specific proposition about retrieval on the tested collections. They do not establish that Embed 5 will outperform competitors on every company's internal records, parsing pipeline or search task.

Cohere's financial-document comparisons are especially relevant to its enterprise positioning. The company says Pro ranked first on its tests of FinanceBench, FinQA and ViDoRe V3 Finance, while Fast placed second on each. Financial filings, reports and spreadsheets often combine dense text with tables and layout that can be lost when documents are converted to plain text. Embed 5 accepts text, images and mixed text-image inputs, and both versions support a 128K-token context window, more than 100 languages and several vector sizes and output formats.

Fast's value proposition also rests on throughput. Cohere reports that it processed documents at an average 2.4 times Pro's throughput in its tests. The company positions Pro for offline indexing and Fast for interactive search, high-volume retrieval and agent workflows, where repeated searches can make latency consequential. Those throughput figures are also company-reported; the announcement does not provide a customer workload or independent deployment result demonstrating the same performance in production.

Embed 5 is available through the Cohere API, Model Vault, Microsoft Foundry and Amazon SageMaker. Cohere's model documentation says it supports single-tenant deployment through Model Vault, giving customers a deployment option beyond shared API access. The company's established focus on enterprise deployments sits behind the model's emphasis on private infrastructure, multilingual documents and retrieval quality.

Gomez's technical history gives the launch a founder-specific throughline: before co-founding Cohere, he worked at Google Brain and co-authored "Attention Is All You Need," the paper that introduced the Transformer architecture. His own biography says he also led researchers at For.ai before Cohere. Embed 5 puts that research legacy to work on a narrower commercial problem: making enterprise data easier to find, at a cost and speed that companies can use in search and agent systems.

Cohere raised $500 million at a $6.8 billion valuation in August 2025, in a round led by Radical Ventures and Inovia Capital, according to the company's funding announcement. Embed 5 does not disclose customer adoption, revenue or independent benchmark validation in its launch material. The immediate bet is measurable in the product design: Cohere is selling retrieval as an operational choice between a higher-scoring index model and a cheaper query model, tied together by a shared vector space.

Reader comments

Conversation for this story loads after sign-in.