Categories / Coding & Development
Embedding API · Pro + FastCohere Embed 5: Pro and Fast for Search and RAG
Choose between retrieval quality and query latency when building enterprise search with Cohere Embed 5. The models produce vectors from retrieval inputs, not generated answers.
Official-source review · Last checked October 3, 2026 · No independent product benchmark
Where Embed fits in retrieval
Cohere Embed 5 converts content and queries into numeric vectors for semantic retrieval. Your application stores document vectors in a search index, compares a query vector against them, and passes selected material to a reader or generation step. It does not itself provide a chatbot, an answer-generating LLM, a reranker, a vector database or a search interface. Embedding concepts.
Use it when you are building search or RAG over enterprise material. If you need a ready-made search interface instead, start with the AI search engine guide.
Pro or Fast: choose the request path
embed-v5.0-pro
Cohere positions Pro for retrieval quality, offline indexing and demanding enterprise corpora.
embed-v5.0-fast
Cohere positions Fast for lower latency and higher throughput: interactive search, agent loops and high-volume queries.
The models share an embedding space. Cohere recommends indexing the corpus with Pro and querying with Fast when latency matters; queries can use either model without rebuilding that index. This is a deployment option to evaluate, not a universally optimal architecture. Cohere release and deployment guidance.
Input and vector contract
The current model table explicitly lists both Pro and Fast with text, image and mixed text/image input, a 128K-token context and dimensions of 256, 512, 768, 1024, 1536 or 2048. Cohere documents multilingual coverage of more than 100 languages.
For document pages, prepare page images or text/image components; do not assume that a raw PDF upload is the same API contract. The Embed v2 reference distinguishes texts, image data URIs and mixed inputs. It supports float, int8 and binary outputs among other encodings. Choose a representation your index can compare and keep dimensions consistent.
The multimodal tutorial demonstrates Pro specifically; its example coverage is narrower than the two-model capability table. We have not executed multimodal requests against either model. Verify the exact model and input format before committing a pipeline.
Access and deployment are separate choices
- Cohere API: use the Embed endpoint with an appropriate API key.
- Microsoft Foundry: a cloud-platform access route announced by Cohere; check the selected model, region and account entitlement.
- Amazon SageMaker: a separate AWS deployment route announced by Cohere; confirm the package and infrastructure requirements.
- Cohere Model Vault: dedicated, Cohere-managed single-tenant inference, rather than shared API access.
The September 30 release lists all four as available. The Model Vault documentation explains its isolation and serving setup. This source review did not provision any of these deployments.
Pricing and production access
Checked October 3, 2026: current exact per-token API pricing was not verified on the pricing page. Consult Cohere pricing for your exact model and access route; launch-announcement rates are not a substitute for a current quote.
The pricing page separately lists Model Vault instance tiers; those are not per-token API prices. Trial keys are rate-limited and cannot be used for production or commercial workloads. Production access requires the relevant account and billing approval. Confirm image billing, rate limits and deployment costs before comparing total cost.
Reported quality versus your own evaluation
Cohere reports retrieval improvements over Embed 4 on enterprise documents and other evaluated corpora. AI Tool Finder did not independently reproduce Cohere's retrieval benchmark. Vendor results and methodology.
Build a small labeled set from the documents you actually retrieve. Compare missed relevant passages, irrelevant matches, query latency and indexing cost. Include scanned pages and each required language. Recheck acceptance thresholds when changing vector dimensions or compression; a smaller index is useful only if retrieval remains good enough for your task.
Continue your implementation
Coding & Development tools · Agent frameworks and retrieval pipelines · All tool guides