AI Tool Review · 2026

Anthropic API Review (2026): Features, Pricing & Verdict

The Anthropic API — the developer platform for the Claude family of models — has grown from the thoughtful challenger of the AI era into one of its two defining foundation-model platforms, and for a large and growing share of serious development work, it’s the one to beat. As with our OpenAI API review, the distinction matters up front: this is the developer API at platform.claude.com, which gives your software programmatic access to Claude models on pay-per-token terms — not the Claude.ai consumer chat app (covered separately at review 0002). Anthropic, founded in 2021 by former OpenAI research leaders Dario and Daniela Amodei, has built its platform on a distinctive thesis: safety-first “constitutional AI,” deep capability in coding and agentic work, and enterprise-grade reliability — and by 2026 the market has emphatically validated it, with the company valued at $380 billion after a $30 billion February raise and revenue in the tens of billions annualised. The models are the draw. Claude leads or co-leads the industry’s coding and agentic benchmarks — Opus 4.8 (May 2026) posts state-of-the-art agentic-coding results, and the entire Claude Code phenomenon is built on these models — while the ladder runs cleanly from Haiku 4.5 ($1/$5 per million tokens) through the workhorse Sonnet 4.6 ($3/$15) and premium Opus 4.8 ($5/$25, a 67% price cut from the old Opus line) to the new Fable 5 frontier tier ($10/$50, launched June 2026). Two structural advantages stand out: a one-million-token context window offered at flat standard rates on current models (rivals surcharge or cap theirs), and cost levers — 90%-off prompt caching, a 50%-off Batch API — that stack to cut real bills dramatically. Add tool use, extended thinking, computer use and MCP, plus availability through AWS Bedrock, Google Vertex AI and Microsoft Foundry, and you have the premier API for coding, agents and long-document work — priced at a premium, and worth it for those workloads.

8.7
Overall Score / 10
The strongest foundation-model API for coding, agents and long-context work — a clean ladder, 1M context at flat rates, deep enterprise trust and excellent cost levers; docked only for premium pricing, a narrower modality surface than OpenAI and a smaller ecosystem
Best for
Developers and companies building coding assistants, autonomous agents, long-document analysis and enterprise AI products — anywhere model reliability, agentic capability and long context matter more than the lowest token price or the broadest modality menu
Platform
The Claude developer API (platform.claude.com) — Messages API with tool use, structured outputs, extended thinking, computer use, prompt caching, Batch processing, Files and MCP support; Python/TypeScript SDKs, workbench and evals. Also on AWS Bedrock, Google Vertex AI and Microsoft Foundry. Not the Claude.ai consumer app
Key differentiator
Best-in-class coding and agentic models (the engine behind Claude Code) plus a 1M-token context window at flat standard rates on current models, wrapped in constitutional-AI safety and the deepest enterprise trust in regulated industries
Pricing
Pay-per-token: Haiku 4.5 $1/$5 per 1M in/out; Sonnet 4.6 $3/$15; Opus 4.8 $5/$25; Fable 5 $10/$50. Prompt caching ~90% off cache hits; Batch API 50% off everything; Opus fast mode 2x for up to 2.5x speed. No ongoing free API tier ($5 trial credits)
Vendor
Anthropic (San Francisco, founded 2021 by Dario & Daniela Amodei) — $380B valuation (Feb 2026), revenue in the tens of billions annualised; SOC 2, no-training-on-API-data by default, data-residency options
Platform notes (2026): a few current specifics worth factoring in. Fable 5 availability episode: Anthropic’s new top-tier Fable 5 (and its trusted-access sibling Mythos 5) was suspended worldwide for 19 days in June 2026 under a US export-control directive following a reported jailbreak; Fable 5 access was restored on 1 July 2026, while Mythos 5 remains limited to approved US organisations — production routers should keep Opus 4.8 or Sonnet 4.6 as fallbacks for frontier-tier traffic. Sonnet 5 introductory pricing: the new Sonnet 5 carries introductory rates of $2/$10 per million tokens through 31 August 2026, after which standard $3/$15 pricing applies — budget for the step-up. And a tokenizer note: Opus 4.7 and later models use a new tokenizer that can generate up to ~35% more tokens for the same text than older models, so benchmark real token counts (not just per-token rates) when migrating. None of these change the platform’s fundamentals, but all three affect planning.

What Is the Anthropic API?

The Anthropic API is the developer platform that provides programmatic access to Anthropic’s Claude models, letting you build text generation, coding, reasoning, agents, vision analysis and document intelligence into your own applications on pay-per-token terms. The distinction from the consumer product mirrors the OpenAI split: Claude.ai is the finished chat application individuals subscribe to (free, Pro at $20/month, Max at $100–200/month), while the API is raw infrastructure — you get an API key, call the Messages API from your backend, and pay for exactly the tokens you use, with no subscription. If you personally want to chat with Claude, subscribe to Claude.ai; if you’re building software that calls Claude, the API is the product, and this review covers that. Anthropic’s origin story shapes the platform’s character. Founded in 2021 by siblings Dario and Daniela Amodei — both former OpenAI research executives — the company was built around AI safety as a first-class engineering discipline, pioneering “constitutional AI,” a training approach in which models follow explicit principles rather than merely pattern-matching from feedback. In practice this translates into models with unusually reliable instruction-following, lower hallucination rates on the Sonnet tier than several peers, and behaviour that enterprises in regulated industries (finance, healthcare, law) have come to trust — a big reason Anthropic’s enterprise business has grown so fast, to a $380 billion valuation and tens of billions in annualised revenue by 2026, with a possible IPO reported for late 2026. The model family is organised as a capability ladder with distinct tiers: Haiku (fastest and cheapest — Haiku 4.5 at $1/$5 per million tokens), Sonnet (the balanced workhorse — Sonnet 4.6 at $3/$15, plus the new Sonnet 5), Opus (the premium tier — Opus 4.8 at $5/$25, launched May 2026), and, as of June 2026, Fable 5 ($10/$50), a new frontier tier above Opus, with a trusted-access variant called Mythos 5 for approved organisations. Around the models sits a modern developer platform: the Messages API, tool use (function calling), structured outputs, extended thinking (visible step-by-step reasoning), vision input, computer use for agents that operate interfaces, prompt caching, a Batch API, Files, and first-class support for MCP (the Model Context Protocol, an open standard Anthropic created that has been widely adopted for connecting AI to tools and data). The API is available first-party and through AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry for teams that need those contractual and regional frameworks. Within our Model Providers & AI Infrastructure category, the Anthropic API is one of the two reference platforms — the coding-and-agents counterpart to the OpenAI API reviewed just before it — and the natural first choice for a large class of serious workloads.

Core Features

The model ladder: coding and agentic leadership from Haiku to Fable

The heart of the Anthropic API is the Claude model family, and its defining strength in 2026 is leadership in exactly the capabilities that matter most for serious production AI: coding, agentic work and long-horizon reasoning. At the top, the ladder has never been stronger. Claude Opus 4.8, released 28 May 2026, is Anthropic’s flagship workhorse for complex work — posting state-of-the-art agentic-coding results (69.2% on SWE-bench Pro, up from 64.3% on Opus 4.7) alongside major gains in long-horizon task coherence and terminal automation — and it arrived at $5/$25 per million tokens, the same rate as its predecessor and a remarkable 67% cheaper than the old Opus 4.x line’s $15/$75, one of the biggest effective price cuts in the API market. Above it sits Claude Fable 5 (June 2026, $10/$50), a new frontier tier delivering the next capability class up for the hardest jobs, with Mythos 5 as a trusted-access variant for approved organisations. The strength of these models in coding is not a niche claim — it’s the foundation of the entire Claude Code phenomenon (Anthropic’s agentic coding tool, itself reportedly a $2.5 billion annualised business), and it’s why Claude models are the default engine inside a large share of the industry’s coding assistants and agent frameworks. Beneath the premium tiers, Claude Sonnet 4.6 ($3/$15) is widely regarded as the best price-to-performance model in the lineup and the sensible default for production — frontier-adjacent quality, a 1M-token context window and adaptive thinking at a mid-tier price — with the new Sonnet 5 arriving at introductory $2/$10 rates (through August 2026). And Claude Haiku 4.5 ($1/$5) anchors the value tier: fast, genuinely capable, and the right default for high-volume classification, extraction and routing. The ladder’s spread — $1 to $10 per million input tokens across four tiers — is narrower than OpenAI’s, but it’s clean and easy to reason about, and the routing pattern is the same: run bulk traffic on Haiku, default production to Sonnet, escalate demanding work to Opus, and reserve Fable for jobs where fewer failed agent loops justify the premium. For teams whose core workloads are coding assistants, autonomous agents or complex multi-step reasoning, this family is, on the 2026 evidence, the strongest available — the reason “just use Claude” became the default advice in developer communities for those tasks.

Long context, extended thinking and the agent toolkit

Anthropic’s second pillar is a set of platform capabilities that compound the models’ strengths — headlined by the industry’s most generous long-context offer and a mature toolkit for building agents. The long-context story is a genuine structural advantage: current Claude models (Fable 5, Opus 4.8/4.7/4.6, Sonnet 5 and Sonnet 4.6) offer a full one-million-token context window at standard pricing — no surcharge, no premium tier — where competitors either cap context lower or apply long-context premiums. A million tokens is roughly 750,000 words: entire codebases, complete contract sets, years of correspondence, or full-length books can be placed in a single request, which transforms what’s possible in document-heavy and codebase-wide work — and it’s precisely the regime (long inputs, agentic loops) where Claude’s coherence advantages show up most. Anthropic pairs this with extended thinking, which lets the model reason step-by-step (visibly, with a budget you control) before answering — invaluable for hard problems — and with the agent toolkit that has made Claude the industry’s preferred agent engine: robust tool use (function calling) with parallel tool execution, structured outputs, computer use (the model operating screens, cursors and terminals for automation), and native support for MCP, the Model Context Protocol. MCP deserves emphasis: Anthropic created it as an open standard for connecting AI models to tools, data sources and services, and it has been adopted across the industry — meaning the connector ecosystem you build for Claude is portable, standards-based and enormous, a quiet but significant moat. Vision input (analysing images, screenshots, charts and documents) is native across the family, and the Files API and code-execution capabilities round out the surface. What Anthropic does not offer matters for fairness: there’s no first-party image-generation model, no video model, and the audio/realtime surface is more limited than OpenAI’s — Anthropic has concentrated deliberately on text, code, reasoning and agents rather than spreading across every modality. For builders whose products need image or audio generation, that means pairing the Anthropic API with a specialist (fal for media, say) or a broader platform. But for the core enterprise workloads of 2026 — agents that write code, read documents, use tools and act reliably over long horizons — Anthropic’s capability set is the most focused and, arguably, the deepest available.

Cost levers, enterprise trust and the multi-cloud footprint

The third pillar is the commercial and enterprise machinery around the models — where Anthropic has quietly built some of the best cost-optimisation tools in the market and the deepest trust with serious buyers. On cost levers, two features stack to cut real bills dramatically. Prompt caching charges cache hits at roughly 10% of the standard input rate — a ~90% discount on repeated context — which is transformative for agents and chat applications that resend large system prompts, tool definitions or document context on every call; for cache-heavy workloads, effective input costs collapse. The Batch API applies a flat 50% discount to both input and output tokens on every Claude model for asynchronous jobs completed within 24 hours, and on current models it also unlocks up to 300K output tokens per request (far beyond synchronous limits) — ideal for bulk generation, classification and document pipelines. Stack the two and well-suited workloads can approach 95% savings versus naive list-price usage; a well-optimised production app on Sonnet with caching typically runs tens of dollars a month at moderate traffic, not hundreds. Additional levers include fast mode on Opus 4.8 (up to 2.5x faster inference at 2x pricing, for latency-critical paths) and, in the other direction, US-only inference at a 1.1x multiplier for workloads with residency requirements. On enterprise trust, Anthropic’s position is arguably the strongest in the industry: constitutional-AI training and a safety-first culture translate into models that regulated industries — fintech, healthcare, law — have adopted at remarkable rates; API data is not used for training by default; SOC 2 compliance, zero-retention and data-residency options are available; and the multi-cloud footprint means Claude is a first-class citizen on AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry, so enterprises can consume it inside their existing cloud contracts, security reviews and regional requirements (regional endpoints carry a ~10% premium over global). The developer experience is polished — clean Python and TypeScript SDKs, a capable workbench, good documentation, an evals framework — and while the surrounding ecosystem remains smaller than OpenAI’s (fewer tutorials, a less universal “compatible-API” gravity), it has grown enormously on the back of Claude Code and MCP. The honest caveat is price positioning: at list rates Anthropic is not the cheapest — budget OpenAI and Google models undercut Haiku, GPT-5.4 undercuts Opus 4.8 at the premium tier, and DeepSeek undercuts everyone — so Anthropic’s economics depend on its levers (caching, Batch) and on the quality argument: fewer failed runs, fewer retries and fewer hallucinations often make the premium model the cheaper total solution. For the workloads Claude leads, that argument usually holds.

Scored Categories

Coding & agentic capability (SWE-bench leadership; Claude Code engine)

9.5

Long-context offer (1M tokens at flat standard rates)

9.3

Model ladder & value (Haiku→Fable; Opus 67% cheaper than old line)

9.0

Safety, reliability & enterprise trust (constitutional AI; regulated industries)

9.0

Agent toolkit (tool use, extended thinking, computer use, MCP)

8.8

Cost levers (90% caching, 50% Batch + 300K output, fast mode)

8.7

Modality breadth & ecosystem size (no image/video gen; smaller than OpenAI)

7.8

List-price competitiveness & availability record (premium; June Fable episode)

7.5

Pricing

Model / item Price (per 1M tokens unless noted) Notes
Claude Fable 5 (frontier) $10 in / $50 out New top tier (June 2026); 1M context at standard rates. Mythos 5 (trusted-access) same price, approved US organisations only
Claude Opus 4.8 (premium) $5 / $25 Flagship workhorse for complex coding, reasoning, agents; 67% cheaper than the old Opus line. Fast mode 2x price for up to 2.5x speed
Claude Sonnet 4.6 / Sonnet 5 $3 / $15 · Sonnet 5 intro $2/$10 Best price-performance; the production default. Sonnet 5 introductory rates through 31 Aug 2026, then $3/$15. 1M context flat
Claude Haiku 4.5 (value) $1 / $5 Fast, capable value tier — the default for high-volume classification, extraction and routing
Prompt caching Cache hits ~10% of input rate ~90% off repeated context; 5-min TTL standard, 1-hour extended available (higher write cost). Transformative for agents/chat
Batch API 50% off input & output All models, 24-hour async window; up to 300K output tokens per request on current models (beta header)
Multi-cloud & modifiers Bedrock / Vertex / Foundry; regional +10% US-only inference 1.1x; data-residency multipliers vary. No ongoing free API tier — new accounts get $5 trial credits; prepaid credits with optional auto-reload
Anthropic’s pricing is a clean three-layer system — base model rates, request-level modifiers, and feature-level charges — and the practical guidance is about using the layers well, because at raw list price Anthropic is not the cheapest API in the market and doesn’t try to be. The base ladder is easy to reason about: $1/$5 (Haiku 4.5) → $3/$15 (Sonnet 4.6) → $5/$25 (Opus 4.8) → $10/$50 (Fable 5), with output uniformly 5x input — so, as everywhere, long outputs are where spend concentrates. Route by task: bulk simple traffic to Haiku, production defaults to Sonnet, hard problems to Opus, and Fable only where fewer failed agent loops justify 2x Opus rates. Then work the levers, because they’re among the best in the market: prompt caching cuts repeated input to ~10% of list — for an agent resending a large system prompt and tool schema every call, this alone can collapse input costs — and the Batch API halves everything for asynchronous work while unlocking 300K-token outputs; stacked, cache-heavy batch workloads approach 95% savings, and a well-optimised Sonnet app at moderate traffic typically lands around $30–100/month. Three planning notes: the 1M-token context at flat rates is a genuine bargain versus surcharging competitors, but a full million-token prompt still costs real money at list rates ($3–10 per uncached call depending on tier) — cache it; the Opus 4.7+ tokenizer can produce up to ~35% more tokens for identical text than older models, so compare real token counts when migrating, not just rates; and there’s no ongoing free API tier ($5 trial credits only), so prototyping beyond that is paid. Comparatively: OpenAI’s GPT-5.4 undercuts Opus 4.8 at list ($2.50/$15 vs $5/$25) and DeepSeek undercuts everyone — but both providers now offer ~90% caching, so effective costs converge for cache-heavy work, and Anthropic’s flat-rate 1M context plus fewer failed runs on agentic tasks frequently makes it the cheaper total solution for the workloads it leads. Verify current rates at claude.com/pricing — Anthropic adjusts the portfolio (though rarely base rates) several times a year.

Strengths

  • Best-in-class coding and agentic models — SWE-bench leadership; the engine behind Claude Code and much of the agent ecosystem
  • 1M-token context window at flat standard rates on current models — no long-context surcharge
  • Clean, easy-to-route model ladder ($1 to $10 per 1M input) with a 67% cheaper premium tier than the old Opus line
  • Excellent cost levers — ~90% prompt caching, 50% Batch API (with 300K output tokens), stacking to ~95% savings on suited workloads
  • Mature agent toolkit — tool use, structured outputs, extended thinking, computer use, vision
  • MCP — the open connector standard Anthropic created, now industry-wide; a portable tooling ecosystem
  • Deepest enterprise trust — constitutional-AI safety, low hallucination rates, adoption across regulated industries
  • No-training-on-API-data by default; SOC 2, zero-retention and data-residency options
  • Multi-cloud availability — AWS Bedrock, Google Vertex AI, Microsoft Foundry
  • Polished DX — clean SDKs, workbench, docs, evals; $380B-backed vendor stability

Weaknesses

  • Not the cheapest at list — GPT-5.4 undercuts Opus 4.8; budget Google/OpenAI models undercut Haiku; DeepSeek undercuts everyone
  • Narrower modality surface — no first-party image or video generation; more limited audio/realtime than OpenAI
  • Smaller ecosystem than OpenAI — fewer tutorials, less universal API-compatibility gravity
  • Output 5x input across the ladder — long generations concentrate cost
  • No ongoing free API tier — $5 trial credits only
  • June 2026 Fable 5 suspension (19 days, export-control directive) showed frontier-tier availability can wobble — keep fallbacks
  • New tokenizer (Opus 4.7+) can add up to ~35% more tokens for the same text — migration cost surprise
  • Fully proprietary — no open weights, no self-hosting

Verdict: 8.7 / 10 — The Coding & Agents Standard

The Anthropic API earns a top-tier 8.7 as the strongest foundation-model platform for the workloads that increasingly define serious AI development — coding, autonomous agents and long-context document work — and as one of the two reference APIs (alongside OpenAI’s) against which the category is judged. Its case is built on genuine, measurable leadership rather than positioning. Claude models lead or co-lead the industry’s coding and agentic benchmarks, power the Claude Code phenomenon and a large share of the agent ecosystem, and pair that capability with the market’s most generous long-context offer — a million tokens at flat standard rates, no surcharge — which together make “build it on Claude” the default advice for coding assistants and agents in 2026. The ladder is clean and well-priced tier-to-tier (with Opus 4.8 arriving 67% cheaper than the old premium line), the cost levers are excellent (90% caching and a 50% Batch API that stack toward 95% savings on suited workloads), the agent toolkit — tool use, extended thinking, computer use, and the industry-adopted MCP standard Anthropic created — is the most focused available, and no vendor commands deeper enterprise trust: constitutional-AI safety, no-training defaults, and multi-cloud availability through Bedrock, Vertex and Foundry have made Claude the choice of regulated industries. What holds it at 8.7 — a hair below our OpenAI API score — are real trade-offs at the margins. Anthropic is a premium product that is rarely the cheapest at list price; its modality surface is deliberately narrower (no first-party image or video generation, limited audio), so multimodal products need a companion platform; its ecosystem, while huge and growing fast, remains smaller than OpenAI’s near-universal gravity; and June 2026’s 19-day Fable 5 suspension was a reminder that frontier-tier availability can be subject to forces beyond any roadmap — keep Opus and Sonnet fallbacks in production routers. None of these dents the core proposition. So the verdict is confident: if you’re building coding tools, agents, document-intelligence products or enterprise AI where reliability and long context matter, the Anthropic API is very likely your best choice — often the cheaper total solution despite premium rates, because better models fail less. If instead you need the broadest multimodal surface under one key, the largest possible ecosystem, or the lowest absolute token price, the OpenAI API, Google’s platform or open-model hosts (Together, Fireworks) are the comparisons to run. The pragmatic 2026 pattern — OpenAI and Anthropic as co-anchors, each routed the workloads it wins — exists precisely because Anthropic earned its half of it.

Frequently Asked Questions

What’s the difference between the Anthropic API and Claude.ai?

They’re two different products built on the same model family, and choosing the right one saves both money and confusion. Claude.ai is the consumer application — the chat interface you use in a browser, desktop app or mobile app, priced by subscription: a free tier (which since mid-2026 runs the capable Sonnet 5 model with usage limits), Claude Pro at $20/month (higher limits, Claude Code in the terminal, projects, integrations), and Claude Max at $100–200/month for power users needing 5x–20x Pro usage. You log in, type, and get answers — no code involved. The Anthropic API is developer infrastructure: programmatic access to the same Claude models via HTTP and SDKs, so your software can call them — you obtain an API key from platform.claude.com, send requests from your application, and pay per token consumed (input and output metered separately), with no monthly subscription; send nothing, owe nothing. Importantly, the two are billed completely separately: a Claude Pro or Max subscription does not include API access, and API credits don’t grant Claude.ai features. The decision rule is the same as for OpenAI’s split: if you’re a person who wants to use AI — to write, code interactively, analyse documents, research — use Claude.ai, where $20/month buys what would cost far more in raw API tokens for a single heavy user. If you’re building a product, feature or automation that calls models programmatically — a support agent, a document pipeline, an AI feature in your app — use the API. One nuance specific to Anthropic: Claude Code, the agentic coding tool, straddles the line — it’s included in Max (and Team Premium) subscriptions or usable via the API at standard token rates, so developers choose based on usage volume (heavy daily use favours the Max subscription; occasional or CI-driven use favours API billing). And note the API has no ongoing free tier — new accounts get $5 in trial credits, after which it’s prepaid pay-as-you-go — whereas Claude.ai’s free tier is genuinely free indefinitely. The kitchen-versus-meal analogy holds: Claude.ai is the finished meal; the API is the kitchen you cook in.

Should I build on the Anthropic API or the OpenAI API?

This is the defining platform choice of 2026, and the honest answer is that both are excellent, the gap is small, and the right pick depends on what you’re building — which is why so many mature teams use both, routed by workload. Choose the Anthropic API when your core workload is coding, agents or long documents. Claude’s leadership on agentic-coding benchmarks is real and sustained (Opus 4.8’s SWE-bench Pro results lead the field), its long-horizon coherence makes agents fail less often — which directly reduces cost, since failed loops are re-billed — and its million-token context at flat rates is unmatched for codebase-wide and document-heavy work (OpenAI’s long context is capable but Anthropic’s flat-rate offer is the more generous deal). Enterprises in regulated industries also tend to favour Anthropic’s safety posture and reliability record. Choose the OpenAI API when you need the broadest capability surface under one key — first-party image generation, a richer audio/realtime stack, built-in hosted agent tools via the Responses API — or when you want the largest ecosystem (more tutorials, more integrations, the de facto compatible-API standard) or the wider price ladder (GPT-4.1 nano’s $0.10/M floor sits below Anthropic’s $1 Haiku floor, and GPT-5.4 undercuts Opus at list). On price, the comparison is subtler than list rates suggest: both now offer ~90% prompt caching and 50% batch discounts, so effective costs converge for optimised workloads, and for agentic tasks Claude’s higher first-pass success rate often makes it the cheaper total solution despite higher per-token rates — measure cost per completed task, not per token. On portability: both have stable, well-documented APIs; Anthropic’s MCP standard means tool/connector investments port broadly, while OpenAI-compatible endpoints mean much basic code ports by changing a base URL — neither locks you in badly at the text-API level, though deeper platform features (OpenAI’s Responses API tools, Anthropic’s computer use) are stickier. The pragmatic recommendation: if you must pick one, pick by primary workload — Anthropic for coding/agents/long-context, OpenAI for multimodal breadth and ecosystem. If you can support two providers (increasingly the norm), anchor each workload on its winner and keep the other as fallback — that portfolio approach is what most sophisticated teams have converged on, and it exists because each platform genuinely wins its lane.

How do I keep Anthropic API costs under control?

Anthropic gives you some of the best cost levers in the API market, and using them deliberately is the difference between a lean bill and an avoidable premium — the playbook has five moves. First and biggest: route by model tier. The ladder spans 10x from Haiku 4.5 ($1/$5) to Fable 5 ($10/$50), and most applications over-provision by defaulting everything to a premium model. Send high-volume simple work — classification, extraction, routing, tagging — to Haiku; make Sonnet 4.6 ($3/$15) your production default (it’s the acknowledged price-performance sweet spot with the same 1M context); escalate to Opus 4.8 only for genuinely hard coding and reasoning; and reserve Fable 5 for frontier jobs where fewer failed agent loops justify 2x Opus rates. Second: exploit prompt caching, which is likely your highest-return single change. Cache hits bill at roughly 10% of the input rate, so any large, stable context — system prompts, tool schemas, reference documents, conversation history — should be structured for caching; for agents that resend a big prefix every call, this collapses input costs by up to 90%, and it’s exactly how teams run 1M-token contexts affordably (pay once to cache the corpus, then query it cheaply). Mind the TTL (5 minutes standard; 1-hour extended caching costs more to write but suits slower loops). Third: batch everything that can wait. The Batch API halves both input and output costs on every model for asynchronous jobs within 24 hours — content pipelines, bulk classification, document processing, evals — and unlocks 300K-token outputs on current models; caching and batch stack, approaching 95% off for suited workloads. Fourth: manage output, because it’s 5x input across the ladder — set max_tokens deliberately, prefer structured/concise output formats, and remember extended-thinking tokens bill as output, so cap thinking budgets to what the task needs. Fifth: watch the migration traps — the Opus 4.7+ tokenizer can generate up to ~35% more tokens for identical text than older models (benchmark real counts, not rates, when upgrading), Sonnet 5’s introductory $2/$10 steps up to $3/$15 after August 2026, and US-only inference or regional endpoints add 1.1x/+10% multipliers. Instrument it all with Anthropic’s usage API and billing alerts. Teams that run this playbook typically land moderate-traffic production apps at $30–100/month on Sonnet; teams that don’t can pay several times more for identical functionality.