Embed | Secure AI Retrieval - Cohere
Skip to content Embed 5 is here: state-of-the-art retrieval - now available in Pro and Fast tiers.

- Embed 5 is here: state-of-the-art retrieval - now available in Pro and Fast tiers.
- Power AI agents that understand your business — retrieving the right data to support reasoning, tool use, and generation across enterprise domains.
- Embed retrieves relevant content in 100+ languages even when queries and source languages don’t match — returning accurate results without any need to identify language or translate.
Embed 5 is here: state-of-the-art retrieval - now available in Pro and Fast tiers. Command: High-performance generative AI models for real-world applications Model Vault provides fully-isolated, performant inference with Saas simplicity How CoreWeave used Cohere North to transform its customer support in 90 days Introducing Embed 5—A new family of frontier embedding models Cohere and Aleph Alpha sign agreement to become the first transatlantic sovereign AI solution The future of work debate has an evidence problem Our state-of-the-art embedding models turn text and images into rich semantic representations that power enterprise search, RAG, and agentic retrieval. Now available in Pro and Fast tiers. Trusted by industry leaders and developers worldwide Transform fragmented data into actionable knowledge From fetching relevant content for Q&A to surfacing task-critical context for agents, Embed enables fast, accurate retrieval across your data. Power AI agents that understand your business — retrieving the right data to support reasoning, tool use, and generation across enterprise domains. The foundation for semantic retrieval in real-world AI systems Built to handle complexity and scale, Embed delivers precise retrieval across noisy, multilingual, and multimodal data. Generate a single embedding for mixed-modality docs containing text, graphs, and tables — simplifying your pipeline and improving accuracy by eliminating data pre-processing. Embed retrieves relevant content in 100+ languages even when queries and source languages don’t match — returning accurate results without any need to identify language or translate. Map visual assets and written content into the same embedding space — making charts, dashboards, and design files just as searchable as any block of text. Embed handles high-context business content with precision — from financial filings to healthcare records — surfacing what’s most relevant, not just what matches the query keywords. Secure, fast, reliable — Everything you need for production-ready retrieval Privately deployable : Run Embed in your virtual private cloud (VPC) or on-premises environment to keep sensitive data secure — or deploy via major cloud services. Efficient at scale : Compress embeddings by up to 96% without sacrificing quality — reducing vector database storage costs and improving performance at the scale of billions of embeddings. Robust in production : Deliver accurate results across noisy, multilingual, and multimodal enterprise data — even on fragmented or domain-specific data. Powered by Command, Compass, Embed, and Rerank, North helps you transform the way you work with secure AI agents, advanced search, and leading generative AI - all in one place. Unlock the potential of your data with an intelligent search and discovery system that doesn't compromise on security. Global organizations choose Embed for search and retrieval “We first stumbled upon Cohere's embedding models because we were using an OpenAI model for embedding. As an experiment, we swapped it out for Cohere’s embedding model and without making any other changes in our code base, we saw all of our metrics around accuracy of response go up.” — Mike Gozzo, Chief Product and Technology Officer, Ada “We first stumbled upon Cohere's embedding models because we were using an OpenAI model for embedding. As an experiment, we swapped it out for Cohere’s embedding model and without making any other changes in our code base, we saw all of our metrics around accuracy of response go up.” — Mike Gozzo, Chief Product and Technology Officer, Ada
Sources
Related stories

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

Predictive Analytics Meets AI
The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.

Mitsuba Squeezes a 27B Vision Model Into 7.3 GB on One GPU
Subtopic Small Models · Vision Language · Quantization Takeaways − Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally
Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.