AI Tool Review · 2026

Cohere Review (2026): Features, Pricing & Verdict

Cohere is the enterprise specialist of the foundation-model market — a company that looked at the consumer AI gold rush and deliberately walked the other way, betting everything on the businesses that can’t, or won’t, send their data to a shared cloud API. Founded in Toronto in 2019 by former Google Brain researchers — including CEO Aidan Gomez, a co-author of the original “Attention Is All You Need” Transformer paper that started this entire era — Cohere has raised over $1.1 billion from NVIDIA, Salesforce Ventures and Cisco, and built a platform whose two signature strengths no larger rival matches. The first is deployment flexibility: Cohere’s models run as a SaaS API, on every major cloud (AWS, Azure, Google Cloud, Oracle, IBM), in your private VPC, fully on-premises, or air-gapped — a spectrum of data sovereignty that is genuinely unique at this capability level, and the reason banking, healthcare, legal and government buyers who cannot use shared APIs keep landing here. The second is the retrieval stack: Cohere’s Embed and Rerank models are widely regarded as the best commercial retrieval tools in the industry — Rerank alone typically cuts RAG hallucination rates by 30–50% by ensuring the language model receives genuinely relevant context — and countless production RAG pipelines use them regardless of whose model does the generating. Around that core sit the Command generation family (from the $0.0375/M Command R7B to the R+ workhorse at $2.50/$10 and the new flagship Command A+ — notably released open-source in May 2026), the North agentic-workplace platform, Compass enterprise search, the new Transcribe speech model, dedicated Model Vault deployments, and the Apache-2.0 North Mini Code coding model. The trade-offs are equally clear: Command trails the frontier flagships on general reasoning, the newest models are contact-sales only, pricing transparency has been slipping, and the enterprise platform carries six-figure commitments. Cohere isn’t trying to be everyone’s AI provider — and for its chosen buyers, that’s exactly the point.

8.1
Overall Score / 10
The enterprise RAG and sovereignty specialist — the industry’s best retrieval stack and unmatched deployment flexibility, held back by a generative frontier gap, growing pricing opacity and enterprise-only gates around its newest models
Best for
Regulated enterprises (banking, healthcare, legal, government) needing private, VPC, on-premises or air-gapped AI; any team building serious RAG or semantic search (Embed + Rerank are best-in-class regardless of your generation model); and organisations wanting commercial-grade models with open-source-grade data control
Platform
Command generation models (R7B, R, R+, A, and the new open-source Command A+ flagship, plus Reasoning/Translate/Vision variants), Embed 4 multimodal embeddings, Rerank 3.5/4 neural rerankers, Transcribe ASR, North AI-workplace/agent platform (with MCP SDK), Compass enterprise search, Model Vault dedicated deployments, fine-tuning, and the Apache-2.0 North Mini Code and Aya open model families
Key differentiator
The industry’s best commercial retrieval stack (Embed + Rerank cut RAG hallucinations 30–50%) combined with deployment anywhere — SaaS, any cloud, VPC, on-prem, air-gapped — a sovereignty spectrum no frontier-class rival offers
Pricing
Pay-as-you-go: Command R7B $0.0375/$0.15 per 1M tokens; Command R $0.15/$0.60; Command R+ $2.50/$10; Embed ~$0.10–0.12/M; Rerank ~$1–2.50 per 1,000 searches. Newest Command A+ family: contact sales. Model Vault dedicated from ~$2,500/month per instance; North/Compass custom. Free trial tier (~1,000 calls/month) for evaluation
Vendor
Cohere (Toronto, founded 2019 by ex-Google Brain researchers incl. Transformer co-author Aidan Gomez) — $1.1B+ raised from NVIDIA, Salesforce Ventures and Cisco; Canadian domicile; B2B-only
Platform notes (2026): three commercial realities to plan around. Pricing transparency is slipping: Cohere’s newest and best models — the Command A+ flagship family, Command A Reasoning, Translate and Vision — are contact-sales only for production keys, the public pricing page renders its current rate tables client-side (so third-party trackers and even careful buyers struggle to verify numbers), and per-minute Transcribe rates aren’t published statically; budget planning here requires a sales conversation in a way it doesn’t at OpenAI or Anthropic. Cloud marketplace availability has shifted: AWS Bedrock’s on-demand Cohere catalogue has been pared back to Embed 4 and Rerank 3.5 — the Command generation family on Bedrock is now provisioned-throughput only, a very different cost shape (roughly $29,000+/month per model unit without commitment) — so verify your intended channel’s current model list and pricing before architecting. The enterprise platform has enterprise gates: North and Compass are demo-and-contract products, reported commitments start in six figures annually, and full deployments take weeks to months — the self-serve API (with its genuine free trial tier of roughly 1,000 calls/month) is a separate, lighter-weight relationship. None of this is unusual for enterprise software; it is unusual for this category, so calibrate expectations accordingly.

What Is Cohere?

Cohere is a business-to-business AI platform company — the only member of this category’s top tier with no consumer product at all — that builds language, retrieval and speech models designed from the ground up for private, secure, enterprise deployment. Its founding pedigree is as strong as any lab’s: established in Toronto in 2019 by former Google Brain researchers, with CEO Aidan Gomez a co-author of the 2017 Transformer paper on which every model in this review series is built, and backing exceeding $1.1 billion from strategic investors including NVIDIA, Salesforce Ventures and Cisco. Its Canadian domicile is itself a quiet selling point — for international buyers wary of US CLOUD Act reach, a non-US-headquartered provider strengthens the privacy posture. The product architecture reflects a single thesis: that the deepest enterprise AI budgets belong to organisations that cannot use shared cloud APIs, and that serving them requires different engineering priorities than serving consumers. Hence the platform’s shape. The Command family provides generation: Command R7B (a remarkable $0.0375/$0.15 per million tokens — among the cheapest production models anywhere), Command R ($0.15/$0.60, the RAG-optimised mid-tier), Command R+ ($2.50/$10, the established workhorse, roughly matching GPT-5.4 on input cost and undercutting it a third on output), Command A with its 256K context for multi-document work, and — as of May 2026 — Command A+, the new mixture-of-experts flagship, which Cohere released open-source, alongside contact-sales Reasoning, Translate and Vision variants. The retrieval models are the crown jewels: Embed 4 (multimodal embeddings at ~$0.10–0.12/M that consistently outperform generic alternatives on enterprise retrieval benchmarks) and Rerank (a per-search-priced neural reranker that has become the industry’s default answer to imprecise vector search). Newer additions extend the surface: Transcribe, a dedicated speech-recognition model (with a best-in-class Arabic variant); North, the flagship agentic AI-workplace platform — agents, intelligent search and task automation with full data isolation, plus an official MCP SDK; Compass enterprise search; North Mini Code, a 30B-parameter MoE coding model released under Apache 2.0 that runs on consumer GPUs; the Aya open multilingual research family; fine-tuning; and Model Vault, dedicated single-tenant model hosting from about $2,500/month per instance. Everything deploys everywhere: SaaS, AWS, Azure, Google Cloud, OCI, IBM Cloud, private VPC, on-premises, air-gapped — with named integrations into Oracle, SAP, Salesforce, Elastic and LivePerson estates. Within our Model Providers & AI Infrastructure category, Cohere is the sovereignty-and-retrieval specialist: rarely the answer to “which model is smartest?”, frequently the answer to “which platform can we actually deploy?”

Core Features

The retrieval stack: Embed and Rerank, the industry’s quiet standard

Cohere’s most important products are arguably not its generation models at all but its retrieval pair — Embed and Rerank — which have become something close to an industry standard for serious RAG and semantic-search systems, adopted even by teams whose generation runs on OpenAI, Anthropic or open models. The problem they solve is the dirty secret of retrieval-augmented generation: vector similarity search is imprecise, and when a RAG pipeline stuffs semi-relevant documents into a model’s context, the model confabulates from bad grounding — most “RAG hallucinations” are really retrieval failures. Embed 4, Cohere’s multimodal embedding model (~$0.10–0.12 per million tokens, English and 100+-language multilingual variants, text and images), attacks the first stage: it consistently outperforms generic alternatives — including OpenAI’s embedding models — on enterprise retrieval benchmarks, meaning the candidate set your vector database returns is better to begin with. Rerank attacks the second stage, and it’s the differentiated one: a neural cross-encoder that takes the user’s query plus the retrieved candidates and re-scores each document for true semantic relevance to that specific question, reordering the list before it reaches the LLM. Priced per search (roughly $1–2.50 per thousand searches depending on tier, with documents over ~510 tokens auto-chunked), it is one of the cheapest interventions in all of AI engineering relative to its impact: production teams report Rerank cutting hallucination rates by 30–50%, because the generator finally receives genuinely relevant context instead of vector-similarity false positives. The full Cohere pipeline — Embed feeds the vector store, Rerank sharpens the top-K, Command generates grounded answers with inline citations — is purpose-built to compose, and at list prices the end-to-end RAG cost frequently undercuts an equivalent GPT-plus-OpenAI-embeddings stack. But the strategic point is that the pieces don’t require each other: a huge share of Rerank’s usage sits inside pipelines that generate with someone else’s model, which makes the retrieval stack both Cohere’s best product line and its most durable — whoever wins the generation wars, retrieval quality remains a separate, purchasable problem, and Cohere owns the best commercial answer to it. For any team building search, knowledge assistants, support deflection or document intelligence, the practical advice is simple and provider-agnostic: benchmark your pipeline with and without Cohere Rerank before spending another dollar on a bigger generation model — the cheaper fix usually wins.

Deploy anywhere: the sovereignty spectrum

Cohere’s second pillar is the deployment story, and it’s the one that wins the seven-figure contracts: no other provider at this capability level lets an enterprise choose its own point on the spectrum from convenient SaaS to fully air-gapped isolation. The spectrum runs in five steps. At the light end, Cohere’s own SaaS API works like any rival’s — keys, per-token billing, a genuine free trial tier (about 1,000 calls a month, generous enough for real evaluation). One step up, cloud marketplaces: Cohere models are available through AWS (Bedrock and SageMaker), Microsoft Azure AI Foundry, Google Cloud, Oracle Cloud Infrastructure, IBM Cloud and beyond — letting enterprises consume Cohere inside existing cloud contracts, security reviews and committed-spend agreements (verify each marketplace’s current catalogue: Bedrock’s on-demand list has narrowed to Embed and Rerank, with Command moving to provisioned throughput). Third, Model Vault: dedicated, single-tenant managed instances of Embed and Rerank from roughly $4–10/hour (~$2,500–6,500/month) — no shared resources, predictable throughput, published pricing. Fourth, private VPC deployment: Cohere’s models running inside the customer’s own cloud perimeter, so data never leaves the organisation’s control. And fifth, the end of the spectrum that defines the company: full on-premises and air-gapped deployment — Cohere’s commercial models running on the customer’s own hardware with no external connectivity at all, an option that OpenAI, Anthropic, Google and xAI simply do not offer for their frontier models, and that otherwise exists only via open-source self-hosting (with all the support and quality trade-offs that implies). This is why Cohere’s positioning is best understood as a hybrid: commercial-grade, professionally supported models with open-source-grade data control — precisely the combination that financial services, healthcare, legal, defence and government buyers require, and the reason Cohere shows up in those industries out of proportion to its consumer mindshare. The North platform extends the same philosophy up the stack: an all-in-one AI workplace — agents, intelligent search, task automation — deployable with full data isolation, integrated with Oracle, SAP, Salesforce, Elastic and LivePerson systems, and speaking MCP via an official SDK, so the agent layer inherits the same sovereignty guarantees as the models beneath it. The honest counterweights: the deep end of the spectrum is contract territory (custom pricing, six-figure commitments, weeks-to-months deployments), and a sovereignty-first roadmap means Cohere invests in isolation and governance features before consumer-visible sparkle. For its buyers, that’s not a bug — it’s the entire pitch.

Command, Transcribe and the open-source turn

The third pillar is the generation-and-beyond model family — solid, well-priced, enterprise-tuned, and lately showing a strategic openness that changes Cohere’s trajectory. The Command family’s design centre is enterprise work rather than benchmark theatre: grounded generation with inline citations, strong tool use for orchestrating business processes (querying CRMs, hitting internal APIs, producing structured outputs), multilingual coverage, and long-document handling (128K contexts standard, 256K on Command A). The ladder prices sensibly: Command R7B at $0.0375/$0.15 per million tokens is one of the cheapest production models in existence — transformative for high-volume classification and extraction — Command R at $0.15/$0.60 is the RAG-optimised value tier, and Command R+ at $2.50/$10 matches GPT-5.4’s input price while undercutting its output by a third. Above them, the 2026 flagship family — Command A+ (a mixture-of-experts frontier model), plus Reasoning, Translate and Vision variants — represents Cohere’s most capable work, with the significant caveat that production access is contact-sales. The honest capability assessment: Command models are excellent at what they’re tuned for — grounded, cited, tool-using enterprise generation — but trail GPT-5.5, Claude Opus and Gemini Pro on open-ended general reasoning and creative work; Cohere isn’t pretending otherwise, and buyers shouldn’t either. Two 2026 additions broaden the surface meaningfully. Transcribe brings dedicated speech-to-text (14 languages, with a best-in-class Arabic variant — a deliberate play for Gulf-region government and enterprise work that fits the sovereignty story), giving voice-heavy enterprises an ASR layer under the same governance umbrella. And the open-source turn is the strategic headline: North Mini Code, Cohere’s first developer-focused coding model — a 30B-parameter MoE with just 3B active parameters, Apache 2.0 licensed, light enough for consumer GPUs — followed by the open-sourcing of the Command A+ flagship itself in May 2026, alongside the long-running Aya multilingual research family. The logic mirrors Mistral’s: open releases build developer mindshare and create an on-ramp to the enterprise platform, while giving sovereignty-minded customers yet another deployment option (run the open weights yourself; buy support, fine-tuning and the platform when it matters). Add accessible fine-tuning for domain adaptation — where a tuned Command R routinely beats a generic frontier model on narrow enterprise vocabulary at a fraction of the cost — and the model story rounds out coherently: not the smartest generalists in the market, but among the best-engineered specialists for the work enterprises actually buy.

Scored Categories

Deployment flexibility & sovereignty (SaaS → VPC → on-prem → air-gapped)

9.5

Retrieval stack (Embed + Rerank best-in-class; 30–50% hallucination cuts)

9.4

Enterprise trust & compliance (regulated-industry fit; Canadian domicile)

8.8

Value at published tiers (R7B $0.0375; R $0.15/$0.60; Embed $0.10)

8.5

Platform breadth (North, Compass, Transcribe, Model Vault, open models)

8.3

Generative capability vs frontier (Command trails GPT-5.5/Opus generalists)

7.3

Ecosystem & self-serve accessibility (small community; enterprise gates)

6.8

Pricing transparency (flagships contact-sales; page opacity; Bedrock shifts)

6.2

Pricing

Product / model Price Notes
Command R7B $0.0375 / $0.15 per 1M tokens One of the cheapest production models anywhere — high-volume classification, extraction, routing
Command R $0.15 / $0.60 per 1M tokens RAG-optimised value tier; 128K context
Command R+ $2.50 / $10 per 1M tokens Established workhorse — matches GPT-5.4 input, ~33% cheaper output; 128K context
Command A+ family (2026 flagship) Contact sales (production) New MoE flagship (open-sourced May 2026) plus Reasoning / Translate / Vision variants; Command A offers 256K context
Embed 4 ~$0.10–0.12 per 1M input tokens Multimodal (text + image), English + 100+-language multilingual; beats generic alternatives on enterprise retrieval benchmarks
Rerank 3.5 / 4 ~$1–2.50 per 1,000 searches Per-search billing (docs >~510 tokens auto-chunk); typically cuts RAG hallucinations 30–50%
Model Vault (dedicated) $4–10/hour per instance (~$2,500–6,500/mo) Single-tenant Embed/Rerank instances; published pricing, predictable throughput
North / Compass / Transcribe / on-prem Custom — contact sales Agent workplace, enterprise search, ASR per-minute rates, VPC/on-prem/air-gapped deployments; six-figure platform commitments reported. Free trial tier ~1,000 calls/month
Cohere’s pricing tells the story of a company mid-pivot from self-serve API to enterprise platform, and buyers should read it in two registers. The published register is genuinely attractive: Command R7B at $0.0375/$0.15 is one of the cheapest production models in existence, Command R at $0.15/$0.60 undercuts every US mid-tier for RAG work, Command R+ at $2.50/$10 beats GPT-5.4 by a third on output, and the retrieval stack is a bargain relative to impact — $0.10/M embeddings plus roughly $2 per thousand reranked searches is frequently the highest-ROI spend in an entire AI budget (a support search handling 100,000 queries a month pays about $200 for Rerank and often saves multiples of that in model-tier downgrades and deflection gains). The free trial tier (~1,000 calls/month, no card pressure) covers honest evaluation, and there are no monthly minimums on pay-as-you-go. The unpublished register is where diligence belongs: the newest Command A+ family is contact-sales for production, the live pricing page renders its current tables client-side (so verify numbers directly, not via trackers), Transcribe’s per-minute rate isn’t published statically, cloud-marketplace pricing is set per marketplace and diverges from cohere.com — with AWS Bedrock’s on-demand catalogue now just Embed and Rerank, and Bedrock Command access moved to provisioned throughput at a very different cost shape (~$29K+/month per model unit) — and the North/Compass platform carries reported six-figure annual commitments with weeks-to-months deployment timelines. Practical guidance: prototype free, run production RAG on the published stack (R or R7B + Embed + Rerank, which is Cohere at its best-value), benchmark Rerank before buying a bigger generator anywhere, and enter platform conversations with procurement expectations calibrated to enterprise software, not API self-serve. Verify current rates at cohere.com/pricing — and expect to talk to a human for anything at the frontier.

Strengths

  • Unique deployment spectrum — SaaS, every major cloud, private VPC, on-premises, fully air-gapped; no frontier-class rival matches it
  • Best-in-class retrieval stack — Embed + Rerank cut RAG hallucinations 30–50%, adopted industry-wide regardless of generation model
  • Deep regulated-industry trust — banking, healthcare, legal, government; full data isolation; Canadian domicile eases CLOUD Act concerns
  • Excellent published value — Command R7B at $0.0375/M among the cheapest production models anywhere
  • Enterprise-tuned generation — grounded answers with inline citations, strong tool use, 256K context on Command A
  • Strategic open-source turn — Command A+ flagship and North Mini Code (Apache 2.0) released openly
  • North agent platform with MCP SDK — sovereignty-grade agents, search and automation; SAP/Salesforce/Oracle integrations
  • New Transcribe ASR (best-in-class Arabic variant) extends the governance umbrella to voice
  • Genuine free trial tier (~1,000 calls/month) and no pay-as-you-go minimums
  • Elite pedigree and stability — Transformer co-author leadership; $1.1B+ from NVIDIA, Salesforce, Cisco

Weaknesses

  • Generative frontier gap — Command trails GPT-5.5, Claude Opus and Gemini Pro on open-ended reasoning and creative work
  • Growing pricing opacity — flagship family contact-sales only; live rate tables render client-side; Transcribe rates unpublished
  • Enterprise gates — North/Compass are demo-and-contract with reported six-figure minimums and weeks-to-months deployments
  • Bedrock on-demand pared back — Command now provisioned-throughput only (~$29K+/month per unit) on AWS
  • Smaller developer ecosystem — fewer tutorials, integrations and community answers than the big three
  • No consumer product means less battle-testing at internet scale and lower mindshare
  • Best models increasingly require a sales relationship — self-serve ceiling is real

Verdict: 8.1 / 10 — The Sovereignty & Retrieval Specialist

Cohere earns a solid 8.1 as the enterprise specialist of the model-provider market — a company that will never top the general-capability leaderboards and has organised itself so that, for its chosen buyers, that fact barely matters. Two of its assets are simply the best available. The retrieval stack — Embed plus Rerank — is the closest thing this industry has to a neutral standard: adopted across pipelines that generate with OpenAI, Anthropic and open models alike, cutting RAG hallucination rates by 30–50% at a cost that makes it the highest-ROI line item in most AI budgets, and durable precisely because retrieval quality is a problem independent of who wins the generation wars. And the deployment spectrum — from SaaS through every major cloud to private VPC, on-premises and air-gapped — is genuinely unique at this capability level, which is why regulated industries that structurally cannot use shared APIs keep signing seven-figure Cohere contracts: it combines commercial-grade models and support with open-source-grade data control, a hybrid position neither the closed US labs nor raw open weights can occupy. The Command family is well-engineered for exactly the work enterprises buy (grounded, cited, tool-using generation at honest prices — R7B at $0.0375/M is a category of one), the 2026 open-source turn (Command A+, North Mini Code) smartly builds the developer on-ramp Cohere historically lacked, and North extends sovereignty guarantees up into the agent layer. What holds the score at 8.1, beneath the general-purpose platforms, is the shape of the specialisation: Command trails the frontier flagships on open-ended reasoning, so Cohere rarely wins deals decided purely on model capability; the ecosystem and self-serve experience trail badly; and — the most concerning 2026 trend — pricing transparency is eroding just as the best models move behind sales conversations, importing enterprise-software friction into a market whose leaders publish every rate. The recommendation is correspondingly precise. If you’re building RAG or semantic search anywhere, benchmark Embed and Rerank this week — provider-agnostic, cheap, and usually decisive. If your organisation needs private, on-premises or air-gapped AI with real support, Cohere is effectively the shortlist. If you’re a generalist buyer choosing one API for everything, the big platforms serve you better. Cohere chose its customers years before they chose it back — and in 2026, that discipline is exactly what its 8.1 is made of.

Frequently Asked Questions

Should I use Cohere or OpenAI/Anthropic for my AI application?

The honest answer is that this is rarely a symmetrical choice — Cohere wins specific, identifiable scenarios decisively and concedes the rest, so the fastest route to a decision is checking which scenario you’re in. Choose Cohere outright in three situations. First, when deployment constraints rule out shared APIs: if your organisation requires private VPC, on-premises or air-gapped AI — common in banking, healthcare, defence, legal and government — Cohere is essentially the only frontier-class commercial option, since OpenAI, Anthropic and Google don’t offer their models for customer-controlled deployment at all, and the alternative (raw open-source self-hosting) trades away professional support, indemnification and managed quality. Second, when your application is retrieval-heavy: for RAG pipelines, semantic search, knowledge assistants and support deflection, Cohere’s Embed and Rerank models are best-in-class regardless of your generation model — many teams run OpenAI or Claude for generation with Cohere handling retrieval, and that hybrid is often the strongest architecture available. Third, when data-residency or jurisdiction matters: Cohere’s Canadian domicile and isolation-first design simplify conversations that US-headquartered providers complicate. Choose OpenAI or Anthropic when raw model capability is the deciding factor — for open-ended reasoning, creative work, cutting-edge coding and agentic tasks, GPT-5.5 and Claude Opus outperform Command, and no deployment flexibility compensates if your product lives on frontier intelligence; when you want the largest ecosystem, richest tooling and most battle-tested infrastructure; or when you’re a startup or individual developer who needs pure self-serve — Cohere’s trial tier is generous, but its best models and platform products increasingly sit behind sales conversations that the big platforms don’t impose at equivalent tiers. The pattern sophisticated teams converge on is unbundling: treat retrieval and generation as separate purchases (Cohere frequently wins the first even when it loses the second), treat sovereignty as a hard filter applied before any benchmark comparison, and let the free tiers do the arguing — an afternoon benchmarking your actual RAG pipeline with and without Cohere Rerank, and your actual prompts on Command R+ versus your incumbent, settles this question with your data rather than anyone’s marketing.

What is Cohere Rerank and why does everyone use it?

Rerank is a neural model that sits between your search retrieval and your language model, re-scoring retrieved documents for true relevance to the user’s specific query — and it has become near-ubiquitous in serious RAG systems because it fixes, cheaply, the failure mode that actually causes most RAG hallucinations. Here’s the mechanism. A standard RAG pipeline embeds a user’s question, pulls the top-K most similar documents from a vector database, and stuffs them into the LLM’s context. The weak link is that vector similarity is a blunt instrument: it retrieves documents that are topically adjacent rather than genuinely relevant — ask “what’s our refund policy for enterprise contracts?” and cosine similarity happily returns the consumer refund policy, a blog post about enterprise sales, and last year’s contract template, because they all live near the query in embedding space. The LLM, handed this plausible-but-wrong context, generates a confident, wrong answer — which gets logged as a “hallucination” when it’s really a retrieval failure. Rerank intervenes at that seam: it’s a cross-encoder that examines the query and each candidate document together (not as independent vectors), scores each for actual semantic relevance to that exact question, and reorders the list so the LLM receives the genuinely relevant material. Because it reads query and document jointly, it catches distinctions embeddings blur — and the measured impact in production systems is a 30–50% reduction in hallucination rates, typically the single largest quality improvement available to a RAG system at any price. And the price is the punchline: Rerank bills per search — roughly $1–2.50 per thousand searches depending on tier — so a knowledge assistant handling 100,000 queries monthly pays a few hundred dollars for what often outperforms upgrading to a generation model costing thousands more. Implementation is a few lines (send query plus candidate documents to the Rerank endpoint; documents over ~510 tokens auto-chunk), it’s generation-model-agnostic — which is why it appears inside pipelines built on GPT, Claude, Gemini and open models alike — and it’s available self-serve, on cloud marketplaces, or as dedicated Model Vault instances for throughput-sensitive deployments. The practical advice this review keeps repeating because it keeps being true: before you buy a bigger LLM to fix your RAG quality, spend an afternoon A/B testing Cohere Rerank on your real queries. It’s the cheapest experiment in AI engineering with the highest hit rate.

Can Cohere really run fully on-premises or air-gapped — and what does that involve?

Yes — genuinely, not as marketing gloss — and it’s the single capability that most separates Cohere from every other provider in this category, so it’s worth understanding what it actually entails. What’s on offer: Cohere’s commercial models (the Command generation family, Embed, Rerank) can be deployed inside your own infrastructure — your private cloud VPC, your on-premises data centre, or a fully air-gapped environment with no external connectivity whatsoever — running on your hardware, under your security perimeter, with no tokens, prompts, documents or telemetry ever leaving your control. This is categorically different from what the big labs sell: OpenAI, Anthropic and Google offer their frontier models only as hosted services (their infrastructure or their cloud partners’), with contractual data protections but never customer-controlled deployment; the only other route to self-hosted AI is open-source weights, which delivers the control but without commercial support, managed updates, indemnification or the quality tier of supported commercial models. Cohere’s hybrid — commercial-grade models with open-source-grade control — is precisely engineered for the buyers in between: banks whose regulators require data never leave approved environments, healthcare systems bound by patient-data rules, defence and government agencies with classified networks, and legal organisations with privilege obligations. What it involves practically: this is the deep end of Cohere’s deployment spectrum, and it’s enterprise business — custom pricing (reported commitments start in six figures annually), a sales-and-solutions engagement rather than a signup form, hardware provisioning on your side (GPU capacity appropriate to your chosen models and throughput), and deployment timelines measured in weeks to months including security review, integration and tuning. Cohere supplies the models, deployment tooling, updates and support; recent open-source releases (Command A+, North Mini Code under Apache 2.0) even give sovereignty-minded teams a lighter self-managed path for appropriate workloads, with the commercial relationship layered on where support and the strongest models matter. The intermediate rungs matter too: many organisations don’t need full air-gapping and land happily on private VPC deployment (data stays inside your cloud account) or dedicated Model Vault instances (single-tenant, published pricing from ~$2,500/month) — so the right conversation with Cohere starts by locating your real requirement on that spectrum rather than defaulting to the most extreme option. If your compliance posture makes shared APIs impossible, this spectrum is Cohere’s entire reason for being — and in 2026 it remains, at frontier-adjacent quality with commercial support, effectively a market of one.