AI Tool Review · 2026

Amazon Bedrock Review (2026): Features, Pricing & Verdict

Amazon Bedrock has quietly become the default control plane for enterprise generative AI — the fully managed AWS service that puts more than 100 foundation models from 18-plus providers behind one unified API, inside the security perimeter, IAM policies and compliance certifications the enterprise already trusts. The catalogue in mid-2026 would have been unthinkable two years ago: the full Anthropic Claude line (still the platform’s most popular family, now including Opus-class models — the latest being the first Bedrock Claude with a 1M-token context window and 128K max output), and — following a significant April 2026 announcement expanding the AWS-OpenAI partnership — frontier OpenAI models including GPT-5.5 and GPT-5.4, plus Codex and the open-weight gpt-oss line, alongside Amazon’s own second-generation Nova family (from $0.035 per million input tokens at the Micro tier up through Premier, with Canvas and Reel for image and video and the Sonic speech-to-speech model that went GA in early 2026), Meta Llama, DeepSeek-R1, Mistral, Cohere, AI21, Stability, Google Gemma, Qwen, NVIDIA Nemotron and more. The platform layers are the real product: batch inference at 50% off, prompt caching cutting costs up to 90%, provisioned throughput for guaranteed capacity, Intelligent Prompt Routing saving up to 30%, AgentCore and Agents for autonomous workflows, Knowledge Bases for RAG, Guardrails at $0.15 per thousand text units (an 80% price cut from launch), Data Automation, model evaluation, fine-tuning and custom model import — all cross-region at source pricing across 44 regions including GovCloud. The honest counterweight is the bill: several CTOs end up spending 1.5–2x their initial estimates — not from hidden fees, but from costs that are difficult to calculate upfront, from the ~$345/month OpenSearch baseline under a “free” Knowledge Base to agent token amplification. This review prices the platform and the traps.

8.5
Overall Score / 10
The enterprise control plane — 100+ models from 18+ providers behind one AWS-secured API with the deepest billing toolkit in the market; docked for cost-predictability traps and AWS-grade complexity
Best for
Enterprises and AWS-native teams that want frontier models — Claude, OpenAI, Nova, Llama, DeepSeek and more — inside their existing VPC, IAM and compliance boundary, with the freedom to switch models by changing an ID
Platform
Fully managed AWS service: unified API over 100+ foundation models; on-demand, batch (−50%), provisioned throughput and prompt caching (−90%) billing; AgentCore, Agents, Knowledge Bases, Guardrails, Flows, Data Automation, fine-tuning/RFT, custom model import; 44 regions incl. GovCloud, cross-region inference at source pricing
Key differentiator
The marketplace-plus-fortress combination — the industry’s broadest first-class model catalogue (including both Anthropic and OpenAI frontiers, a 2026 first) delivered inside AWS’s security, compliance and networking perimeter, at price parity with going direct
Pricing
Per-token, no flat fee: Nova Micro from $0.035/1M input to frontier models at $18+/1M output; Claude Sonnet-class ~$3/$15 (parity with direct); batch 50% off; caching up to 90% off; Guardrails $0.15/1K text units; Knowledge Bases free but underlying vector store billed (OpenSearch ~$345/mo baseline; S3 Vectors up to 90% cheaper). Typical bills: ~$100/mo prototypes to $5,000+/mo agentic production
Vendor
Amazon Web Services — the world’s largest cloud, with Bedrock as its generative-AI centrepiece and the Nova model family as its first-party line
Platform notes (2026): four budget realities before you build. The estimate gap is documented: real Bedrock bills commonly land 1.5–2x initial projections — not hidden fees, but hard-to-model costs: output tokens priced 3–5x input, agent workflows amplifying token counts across steps, and adjacent AWS services (CloudWatch, data transfer at $0.09/GB out to internet) accruing beside the model spend. The Knowledge Base trap: Knowledge Bases carry no separate fee, but the default vector store — OpenSearch Serverless — has a 2-OCU minimum baseline of roughly $345/month even at zero query traffic; S3 Vectors (December 2025) is up to 90% cheaper and the 2026 default recommendation for most RAG workloads. Model availability varies by region: the newest frontier models land in us-east-1 first and roll out unevenly across the 44 regions — verify your region’s catalogue before architecting, though cross-region inference (source-region pricing, no routing surcharge) covers many gaps. Agents bill by what they trigger: Bedrock Agents and AgentCore have no separate platform fee — you pay for the underlying model calls, retrievals and guardrail evaluations they generate, which is exactly where multi-step reasoning quietly multiplies token spend.

What Is Amazon Bedrock?

Amazon Bedrock is AWS’s fully managed foundation-model service — the layer that turns the world’s largest cloud into a generative-AI platform — and its defining idea has held steady since launch while everything around it scaled: Bedrock is not a model, it’s a control plane. The model you call is a configuration decision — change a model ID and you’ve switched providers — while the security, networking, observability and billing layer stays constant and stays AWS. What has transformed is the catalogue behind that API. From a curated handful at launch, Bedrock in mid-2026 fronts more than 100 foundation models from 18-plus providers, and the roster reads like a peace treaty: Anthropic’s full Claude line (the platform’s most popular family, through the Claude 4 generation, with the newest Opus-class model bringing a 1M-token context and 128K output to Bedrock for the first time); OpenAI’s frontier models — GPT-5.5, GPT-5.4, Codex and the open-weight gpt-oss line — added through the April 2026 partnership expansion that ended the era of OpenAI-means-Azure exclusivity; Amazon’s own Nova 2 family spanning Micro/Lite/Pro/Premier text tiers, Canvas (image), Reel (video) and Sonic (speech-to-speech, GA early 2026); plus Meta Llama, DeepSeek-R1, Mistral, Cohere, AI21, Stability AI, Google Gemma, Qwen, Writer, Luma AI, TwelveLabs and NVIDIA Nemotron, with a Marketplace extending into specialised and emerging models beyond the core catalogue. Around the models, the platform stack is what enterprises actually buy: Bedrock Agents and the newer AgentCore for autonomous multi-step workflows; Knowledge Bases for managed RAG (bring your data, Bedrock handles chunking, embedding and retrieval); Guardrails for content filtering and hallucination checks at $0.15 per thousand text units — 80% cheaper than at launch; Flows for visual workflow orchestration; Data Automation for turning unstructured documents, audio and video into structured output; model evaluation (automatic, or human at $0.21 per task); fine-tuning, reinforcement fine-tuning and distillation; custom model import (bring your own weights, pay only per model-copy serving time); and Intelligent Prompt Routing, which routes each request between model tiers and can cut costs up to 30% without accuracy loss. Billing comes in modes that reward planning: on-demand per-token (the default), batch at 50% off for asynchronous jobs, provisioned throughput for reserved capacity on 1- or 6-month commitments, and prompt caching that can cut input costs up to 90% on repetitive-context workloads. All of it runs inside the AWS trust perimeter — VPC isolation, IAM, CloudTrail, the compliance certification stack, 44 regions including GovCloud, and cross-region inference billed at source-region rates with no routing surcharge. Within our Model Providers & AI Infrastructure category, Bedrock is the aggregation thesis at enterprise scale: where every other review in this series asks “is this provider’s model worth adopting,” Bedrock asks “why adopt a provider at all when you can adopt them all behind one door” — and for a large share of the enterprise market, that question has already been answered.

Core Features

The catalogue: one API, every frontier

Bedrock’s first-order feature is breadth with neutrality, and 2026 is the year the breadth became genuinely complete. The strategic milestone was April’s OpenAI partnership expansion: for the first time, an enterprise can call Anthropic’s Claude Opus-class models and OpenAI’s GPT-5.5 through the same API, the same IAM policies, the same VPC endpoints and the same bill — which converts the industry’s defining rivalry into a dropdown menu. The practical consequences are larger than they sound. Model migration collapses from a re-architecture project to a configuration change: the unified Converse API means switching from Claude to Nova to Llama is an ID swap, and Intelligent Prompt Routing automates the choice per-request (routing simple queries to cheap models, complex ones to frontier models, with documented savings up to 30% at equivalent accuracy). Vendor negotiation leverage inverts: teams that once locked into one provider’s pricing now benchmark continuously across the catalogue, and Bedrock’s price parity with going direct (Claude Sonnet-class at the same ~$3/$15 you’d pay Anthropic; in-region OpenAI inference at parity with OpenAI’s data-residency tier) means aggregation costs nothing at the token level. Coverage extends across modalities — text and code from every major family, DeepSeek-R1 for open reasoning at high-volume prices, Stability and Nova Canvas for image, Nova Reel and Luma for video, Nova Sonic for speech-to-speech, TwelveLabs for video understanding, plus embeddings and rerank models for the RAG stack — and across trust tiers, from open weights (gpt-oss, Llama, Gemma, Qwen) you could later self-host, to frontier closed models under enterprise terms. The first-party Nova 2 family anchors the value floor: Micro at $0.035 per million input tokens is one of the cheapest capable models anywhere — the routing/classification/extraction workhorse — with Lite around $0.06, Pro at $0.80/$3.20 handling mid-tier work, and Premier at the top; because they’re first-party, Nova models are consistently the platform’s price leaders, and for the large fraction of enterprise workloads that are structured extraction and classification rather than frontier reasoning, they carry the load at near-negligible cost. The honest limits of the catalogue thesis: the newest releases still reach their home platforms first (a new Claude or GPT typically lands on Anthropic/OpenAI direct days-to-weeks before Bedrock, and regional rollout across 44 regions is uneven), a handful of notable providers remain absent or partial, and the sheer menu — 79-plus models with published pricing spanning $0.04 to $18.80 per million input tokens — shifts the burden of model selection onto the buyer, which is a real cost the single-provider path doesn’t carry.

The enterprise stack: agents, RAG, guardrails and governance

The catalogue gets the headlines; the platform services around it are why enterprises standardise on Bedrock rather than a thin router. The agent layer leads in 2026: Bedrock Agents configure autonomous assistants that connect to data sources, maintain session memory, call tools and APIs, and decompose tasks into steps — with AgentCore, the 2026 addition, providing the production-grade runtime (identity, memory, observability, tool gateways) that turns agent prototypes into governed deployments; neither carries a separate platform fee, with billing flowing from the model calls, retrievals and guardrail checks the agent triggers — an elegant structure whose flip side is that multi-step reasoning multiplies token consumption in ways that dominate real agentic bills. Knowledge Bases deliver managed RAG end-to-end — point them at your documents and Bedrock handles chunking, embedding, vector storage and retrieval, with rerank models (billed per query, 100 chunks per query unit) sharpening relevance — and the critical 2026 buying note is the vector-store choice underneath: OpenSearch Serverless, the historical default, carries a ~$345/month two-OCU floor even at zero traffic, while S3 Vectors (launched December 2025) runs up to 90% cheaper, supports trillions of vectors at sub-second latency, and is now the right default for most workloads. Guardrails provide model-agnostic safety — content filters, denied topics, PII redaction, contextual-grounding checks against hallucination — at $0.15 per thousand text units after an 80% price cut, with no charge for blocked requests; because they’re a platform layer, one policy applies identically across Claude, GPT, Nova and Llama, which is the governance property compliance teams actually want. The remaining services fill out the lifecycle: Flows for visual orchestration of models, agents and AWS services; Data Automation for industrialised document/audio/video-to-structured-data processing (the intelligent-document-processing workhorse); model evaluation with algorithmic scores free beyond inference and human evaluation at $0.21 per task; fine-tuning and the newer reinforcement fine-tuning (Bedrock automates the RFT loop against your reward function, billed hourly, with the tuned model served on-demand after); distillation for shrinking frontier behaviour into cheap models; and custom model import, which serves your own fine-tuned open weights inside Bedrock’s managed environment billed per model-copy in five-minute increments — no charge for the import itself. Wrapped around everything is the property no standalone provider can replicate: the AWS operational fabric — IAM-scoped access per model per team, CloudTrail audit on every invocation, CloudWatch token/latency/spend metrics, VPC endpoints keeping traffic off the public internet, and the compliance certification stack (HIPAA, SOC, FedRAMP-adjacent GovCloud) that lets regulated industries deploy generative AI without a new security review per vendor. This stack is Bedrock’s real moat: models are increasingly commodities; governed, observable, multi-model production infrastructure is not.

Billing modes, optimisation levers and the true cost of ownership

Bedrock’s billing is simultaneously the most flexible in this category and the most demanding of operator skill — the gap between a tuned deployment and a naive one routinely exceeds the gap between providers. The modes: on-demand (per-token, no commitment, the default — with cross-region inference at source-region pricing giving capacity resilience for free); batch at a flat 50% discount for asynchronous jobs (the single easiest saving for any workload that tolerates latency — evaluations, backfills, bulk document processing); provisioned throughput, which works like reserved instances — model units guaranteeing throughput, billed hourly on 1- or 6-month terms, economical above roughly 60–70% utilisation and a money-loser below; and prompt caching, which discounts repeated context by up to 90% and is the highest-leverage optimisation for agent and RAG workloads whose prompts share large static prefixes — documented as dramatically underused. Layer the routing tools on top — Intelligent Prompt Routing between model tiers (up to 30% savings at maintained accuracy), and simple model right-sizing across the catalogue (the difference between reflexively calling a frontier model and routing classification to Nova Micro is literally two orders of magnitude: $0.035 versus multi-dollar input rates) — and disciplined teams can drive unit costs down relentlessly. Now the honest ledger of where undisciplined budgets die, because the documented pattern is consistent: real bills land 1.5–2x initial estimates, and the causes are structural, not hidden. Output-token asymmetry — across nearly every model, output costs 3–5x input (Claude Sonnet-class at $3/$15 being the canonical example), so chat-heavy and verbose-reasoning workloads cost far more than input-based estimates suggest. Agent amplification — a single user request through a multi-step agent generates many model calls, retrievals and guardrail evaluations, each billed; agentic production is where “$100/month prototype” becomes “$5,000/month platform.” Adjacent services — the OpenSearch floor under Knowledge Bases, CloudWatch metrics and logs, data transfer ($0.09/GB out to internet; free within-region with VPC endpoints), evaluation-judge tokens billed at standard rates — accrue beside the Bedrock line item and outside most estimates. And customisation carries its own meter: fine-tuning and RFT bill training hours, distillation bills the teacher model’s synthetic-data generation at on-demand rates, custom model copies bill serving minutes whether traffic arrives or not. None of this is predatory — every rate is published — but Bedrock is a platform whose economics reward engineering attention and punish set-and-forget, and the right mental model is the one AWS veterans already hold: the bill is a design artefact.

Scored Categories

Model catalogue breadth (100+ models, 18+ providers — both Claude and GPT frontiers)

9.6

Enterprise security & compliance (VPC, IAM, CloudTrail, 44 regions, GovCloud)

9.4

Vendor neutrality & switching freedom (model = config change; parity pricing)

9.0

Platform capabilities (AgentCore, Knowledge Bases, Guardrails, Data Automation, Flows)

8.8

Billing flexibility (on-demand, batch −50%, provisioned, caching −90%, routing −30%)

8.6

Price competitiveness (parity with direct; Nova floor at $0.035/1M input)

8.2

Developer experience (unified Converse API vs AWS console sprawl and IAM learning curve)

7.6

Cost predictability (documented 1.5–2x estimate overruns; adjacent-service accrual)

6.8

Pricing

Item Price Notes
Amazon Nova Micro From $0.035 / 1M input tokens The platform’s value floor — routing, classification, extraction. Nova Lite ~$0.06; Nova Pro $0.80/$3.20; Premier above
Anthropic Claude (Sonnet-class) ~$3 / $15 per 1M tokens Parity with Anthropic direct; full Claude line through Opus-class with 1M context on Bedrock. Output = 5x input — the asymmetry that drives chat-heavy bills
OpenAI frontier models (GPT-5.5 / 5.4, Codex, gpt-oss) In-region at parity with OpenAI’s data-residency tier April 2026 partnership; global cross-region pricing rolling out
Catalogue range $0.04 – $18.80 / 1M input across 79+ priced models DeepSeek-R1, Llama, Mistral, Cohere, AI21, Gemma, Qwen, Nemotron and more; per-image / per-minute billing for media models
Batch inference 50% off on-demand Asynchronous jobs to S3 — the easiest structural saving for latency-tolerant work
Prompt caching Up to 90% off cached input Highest-leverage lever for agents and RAG with shared context prefixes; widely underused
Provisioned throughput Hourly per model unit, 1- or 6-month terms Reserved-capacity economics — pays above ~60–70% utilisation, costs below it
Guardrails $0.15 / 1K text units 80% cheaper than launch; no charge for blocked requests; same policy across all models
Knowledge Bases No platform fee — components billed ⚠ OpenSearch Serverless default: ~$345/mo two-OCU floor at zero traffic. S3 Vectors (Dec 2025): up to 90% cheaper — the 2026 default choice
Agents / AgentCore No platform fee — pay what they trigger Model calls, retrievals, guardrail checks per step; token amplification is the real agentic cost
Evaluation / fine-tuning / import Various meters Human eval $0.21/task; RFT and fine-tuning billed hourly; distillation bills teacher tokens; custom model copies billed per 5-min serving increments
Typical monthly bills ~$100 (prototype) → $5,000+ (agentic production) Documented pattern: real bills 1.5–2x naive estimates — budget the adjacent services
Bedrock budgeting has a documented failure mode and a documented cure. The failure: estimating from the per-token table alone — the real bill adds output asymmetry (3–5x input), agent-step amplification, the vector-store floor, CloudWatch and data transfer, and lands 1.5–2x the projection. The cure, in order of leverage: cache aggressively (up to 90% off repeated context — the single biggest saving for agents and RAG); batch everything latency-tolerant (flat 50% off); right-size ruthlessly (Nova Micro at $0.035 handles the classification and extraction that doesn’t need a $3-input frontier model — a 100x spread you control with a model ID); choose S3 Vectors over OpenSearch for Knowledge Bases unless you need OpenSearch’s features (up to 90% cheaper, no $345 floor); use Intelligent Prompt Routing for mixed traffic (up to 30% off at maintained accuracy); and commit only above ~70% utilisation — provisioned throughput below that line pays for idle capacity. Verify current rates at aws.amazon.com/bedrock/pricing before budgeting: the catalogue and its prices change monthly.

Strengths

  • The industry’s broadest first-class catalogue — 100+ models, 18+ providers, and uniquely both Claude and GPT frontiers behind one API
  • AWS-grade security and governance — VPC, IAM, CloudTrail, compliance stack, 44 regions, GovCloud
  • Model switching as configuration — an ID change, not a migration; parity pricing means aggregation is free at the token level
  • Deepest billing toolkit in the category — batch −50%, caching −90%, routing −30%, provisioned capacity
  • Complete enterprise stack: AgentCore, Knowledge Bases, Guardrails (80% cheaper), Data Automation, evaluation, fine-tuning, custom import
  • Nova family anchors a genuine value floor from $0.035/1M input
  • Cross-region inference at source pricing — capacity resilience with no surcharge
  • No flat fee — pure usage-based entry from prototype scale

Weaknesses

  • Documented cost-predictability problem — real bills commonly 1.5–2x estimates from adjacent services and token amplification
  • The Knowledge Base vector-store trap — ~$345/mo OpenSearch floor at zero traffic unless you choose S3 Vectors
  • AWS complexity tax — IAM, console sprawl and service interdependencies raise the learning curve well above single-provider APIs
  • Newest frontier releases reach home platforms first; regional catalogue rollout is uneven
  • Output tokens at 3–5x input rates punish verbose and chat-heavy workloads
  • Provisioned throughput punishes over-commitment below ~60–70% utilisation
  • Model-selection burden shifts to the buyer — 79+ priced models is a menu, not a recommendation
  • AWS gravity — neutrality among models, but the platform itself deepens cloud lock-in

Verdict: 8.5 / 10 — The Enterprise Control Plane

Amazon Bedrock earns an 8.5 as the definitive answer to a question most of this category’s reviews can’t ask: what if you didn’t have to choose? The 2026 catalogue — Claude and GPT frontiers together for the first time, Nova’s $0.035 floor, DeepSeek and Llama and Mistral and a hundred more behind one API at parity pricing — makes model choice a reversible configuration decision, and the AWS fabric around it (VPC, IAM, Guardrails spanning every model identically, 44 regions, GovCloud) is the governance layer no standalone provider can match and every regulated enterprise eventually requires. The billing toolkit is genuinely best-in-class — batch, caching, routing and provisioned capacity compose into unit costs disciplined teams drive relentlessly down. What holds it at 8.5 rather than higher is that the platform’s power is priced in operator skill: the documented 1.5–2x estimate overruns, the OpenSearch floor lurking under Knowledge Bases, agent token amplification, output-token asymmetry and the AWS learning curve mean Bedrock rewards teams with cloud engineering maturity and quietly taxes teams without it — the same workload costs dramatically different amounts depending on who architected it. The buying logic: if you’re an enterprise or AWS-native team, Bedrock is the default — start on-demand with Nova for the cheap tiers and Claude or GPT for the frontier work, cache and batch from day one, choose S3 Vectors for RAG, and treat Intelligent Prompt Routing as free money. If you’re a startup optimising raw price-performance on one model family, going direct (or via the open-weights hosts in this series) is simpler and occasionally faster to the newest releases. And if you’re anyone building agents for production, Bedrock’s AgentCore-plus-Guardrails-plus-CloudTrail combination is currently the most credible governed path from demo to deployment — just budget it like the multi-service AWS architecture it is, because that’s exactly what you’re buying.

Frequently Asked Questions

Should I use Bedrock or go direct to Anthropic / OpenAI?

The decision reduces to three questions — where your infrastructure lives, how many models you’ll use, and how much governance you need — and the price question, surprisingly, mostly drops out. On price: Bedrock serves the major families at parity with going direct (Claude Sonnet-class at the same ~$3/$15; in-region OpenAI inference at parity with OpenAI’s data-residency tier), so aggregation itself costs nothing per token — the differences live in the billing toolkit (Bedrock’s batch −50%, caching −90% and provisioned modes match or exceed what the direct APIs offer) and in the adjacent-service costs unique to Bedrock deployments. Choose Bedrock when: your workloads already live on AWS (VPC endpoints keep traffic private, IAM scopes access per team per model, CloudTrail audits every call — properties that would take months to replicate around a direct API); you want multi-model freedom (the ability to A/B Claude against GPT against Nova by changing an ID, route intelligently between tiers, and renegotiate with leverage — impossible in a single-provider relationship); you’re in a regulated industry (the compliance certifications, GovCloud availability and model-agnostic Guardrails layer are the fastest path through a security review); or you’re building production agents (AgentCore’s managed runtime plus platform-level observability is the most governed agent path currently shipping). Choose direct when: you’re a startup or small team on a single model family, where the provider’s native API is simpler, its SDK ergonomics are better, and its newest releases arrive first — the home-platform lag on Bedrock is typically days to weeks for frontier launches, which matters if you live on the bleeding edge; you need provider-native features that lag on Bedrock (each provider’s most experimental capabilities ship to its own API first); or your organisation has no AWS footprint, in which case Bedrock’s IAM and networking prerequisites are a tax with no offsetting benefit. The hybrid pattern — increasingly the norm among sophisticated teams — uses both: direct APIs for development speed and newest-model access, Bedrock for production deployment where governance, multi-model routing and the billing toolkit earn their complexity. And one honest asymmetry deserves the last word: migrating from direct to Bedrock later is straightforward (the models are the same); migrating a production system’s worth of IAM policies, Guardrails and Knowledge Bases off Bedrock is not — the neutrality among models is real, but the platform itself is AWS gravity, and you should walk in knowing it.

What actually drives Bedrock bills over budget — and how do I prevent it?

The documented pattern is that real Bedrock bills land 1.5–2x initial estimates, and the causes are structural and preventable — five mechanisms account for nearly all of it. One: output-token asymmetry. Estimates built on headline per-token rates usually anchor on input pricing, but across nearly every model output costs 3–5x input ($3/$15 on Claude Sonnet-class being canonical), so chat applications, verbose reasoning modes and long-form generation cost multiples of the naive projection; prevention is estimating from realistic input:output ratios (measure a pilot; chat commonly runs 1:2 to 1:4) and setting max-token limits on responses. Two: agent amplification. A single user request through a multi-step agent triggers many model invocations — planning, tool calls, retrievals, reflection, guardrail checks per step — so agentic workloads consume tokens at a multiple of chat workloads; prevention is instrumenting per-request step counts early (CloudWatch metrics expose this), capping agent iterations, and routing intermediate steps to cheap models (Nova Micro planning, frontier-model finals). Three: the vector-store floor. Knowledge Bases are “free” but their default OpenSearch Serverless backend bills a two-OCU minimum (~$345/month) at zero traffic — the classic surprise line item on prototype bills; prevention is choosing S3 Vectors (December 2025, up to 90% cheaper, trillions of vectors, sub-second latency) unless you specifically need OpenSearch capabilities. Four: adjacent AWS accrual. CloudWatch logs and metrics, data transfer ($0.09/GB out to internet; $0.02/GB cross-region), NAT gateway charges when VPC endpoints aren’t used, evaluation-judge tokens billed at standard rates — each small, collectively material; prevention is same-region architecture, VPC endpoints (which also eliminate NAT charges), and tagging Bedrock resources for cost-explorer attribution from day one. Five: customisation meters. Distillation bills the teacher model’s synthetic-data generation at on-demand rates, fine-tuning and RFT bill training hours, and imported custom models bill per-copy serving minutes whether or not traffic arrives; prevention is costing the full training-plus-serving lifecycle before committing, and scaling custom model copies to zero in dev environments. Then flip from defence to offence, because Bedrock’s optimisation levers are equally structural: prompt caching (up to 90% off repeated context — transformative for agents and RAG, and documented as the most underused lever on the platform), batch mode (flat 50% off anything asynchronous), model right-sizing (the $0.035-to-$18.80 catalogue spread means routing decisions dominate rate negotiations), Intelligent Prompt Routing (up to 30% automatic savings on mixed traffic), and provisioned throughput only above ~70% sustained utilisation. Teams that operationalise this list — measure, cache, batch, route, right-size — reliably land under their original estimates; the platform punishes only the set-and-forget.

What are Amazon’s Nova models, and are they good enough to matter?

Nova is Amazon’s first-party foundation-model family — second-generation as of 2026 — and the correct frame is not “is Nova the best model” (it isn’t, and AWS knows it) but “what fraction of enterprise workloads does Nova handle at prices that make frontier models look absurd” — to which the answer is: a large one. The lineup: four text tiers — Micro (from $0.035 per million input tokens, one of the cheapest capable models on any platform), Lite (~$0.06 input, multimodal), Pro ($0.80/$3.20, the mid-tier workhorse) and Premier (the reasoning flagship) — plus Canvas for image generation, Reel for video, and Sonic, the speech-to-speech model that went GA in early 2026 and handles real-time voice conversation as a single model rather than an ASR-LLM-TTS pipeline. Because they’re first-party, Nova models are structurally the platform’s price leaders — AWS prices them to win the volume tiers — and they integrate deepest with the platform features (Intelligent Prompt Routing’s documented pairs include Nova Pro/Lite routing; custom Nova fine-tunes serve at base-model rates). Where Nova genuinely earns its place: the unglamorous majority of enterprise AI — classification, entity and structured extraction, summarisation at volume, routing and triage, document processing through Data Automation, embeddings-adjacent work — where a task either succeeds or fails and “frontier-model prose quality” buys nothing; at $0.035 input, Nova Micro processes a million routing decisions for what a frontier model charges for a few thousand, and in the standard architecture (cheap model triages, expensive model escalates) Nova is the triage layer that makes the economics work. Where Nova honestly doesn’t compete: frontier reasoning, complex code generation, subtle long-form writing and the hardest agentic work — the territory where Claude, GPT-5.5 and the top open models live, and where Bedrock’s own catalogue design (Nova beside the frontiers, one API) implicitly concedes the point; Premier narrows the gap but doesn’t close it, and nobody serious builds their escalation tier on it today. The verdict embedded in the design: Bedrock’s aggregation thesis works precisely because Nova exists — the platform can promise “right-size every workload” only because it owns a price-floor family it can serve at cost — and the practical guidance for any Bedrock deployment is to benchmark Nova Micro/Lite on your high-volume, low-complexity traffic first, route what passes, and spend the savings on frontier tokens where they actually change outcomes. Used that way, Nova is less a model choice than a cost-architecture primitive — and one of the quiet reasons Bedrock deployments that follow the playbook come in under budget.