Elevata

Article

Claude Sonnet 5 on Amazon Bedrock: Cost and Model Choice

Paulo Frugis
View profilePublished June 30, 2026Updated September 10, 20269 min read

Anthropic published Claude Sonnet 5 on June 30, 2026, and AWS announced availability on Amazon Bedrock and Claude Platform on AWS. For AWS teams, the adoption decision is whether Sonnet 5 meets the workload’s quality, latency, and cost requirements, and which tasks still justify Opus 5.

The short answer: start testing Sonnet 5 for high-volume, repeatable, reviewable workloads. Route to Opus 5 when one more correct answer materially changes the outcome or when failure is expensive to fix. Do not switch production from a public benchmark alone: use your prompts, tools, data, latency, review burden, and real cost.

Claude Sonnet 5 at a glance

QuestionPractical answer
What is it?Anthropic's most capable Sonnet model so far, aimed at coding, agents, and professional work at scale.
Public announcementAnthropic’s announcement and the current AWS Bedrock model card both list June 30, 2026.
AWS availabilityAmazon Bedrock and Claude Platform on AWS. On Bedrock, paths include bedrock-runtime and bedrock-mantle.
Main Bedrock IDsRuntime: us.anthropic.claude-sonnet-5, eu.anthropic.claude-sonnet-5, au.anthropic.claude-sonnet-5, or global.anthropic.claude-sonnet-5. Mantle: anthropic.claude-sonnet-5.
Context and outputAWS's model card lists a 1M-token context window and 128K-token maximum output. Anthropic's docs say the 1M window is both the default and maximum.
PricingStandard global: $2 input / $10 output per million tokens. Supported commercial geographic or in-Region routing: $2.20 / $11. GovCloud: $2.40 / $12.
First workload to testHigh-volume coding agents, review, documentation, support, analysis, and automation where Sonnet 4.6 was close but Opus was too expensive.

Sources for fast-changing facts: the AWS Bedrock model card, Amazon Bedrock pricing, and Anthropic's Sonnet 5 model notes.

What the benchmarks should and should not change

Anthropic positions Sonnet 5 as a strong improvement over Sonnet 4.6 and close to Opus 4.8 in several agent and professional-work evaluations. That is a useful signal, but it is not a deployment plan. Benchmark curves compress different concerns into one score: success rate, latency, reasoning effort, token use, prompt stability, and error recovery.

Use the benchmarks this way:

  • For default model selection: Sonnet 5 is now a credible first model for many production Bedrock workloads that previously needed an Opus trial.
  • For premium-model justification: reserve Opus for the ambiguous, high-stakes, or failure-expensive tasks where a small accuracy gain changes the business result.
  • For effort controls: do not compare only top benchmark settings. Higher reasoning effort can improve quality but also changes latency and cost.
  • For migration: re-run your real eval set. A public benchmark will not reveal your tool-call failure modes, schema breakage, retrieval drift, or human-review burden.

This is the same reason our Opus 4.8 benchmark guide argues against treating public leaderboards as procurement decisions. The ranking tells you where to investigate. It does not tell you what to ship.

What Artificial Analysis found about cost per task

Artificial Analysis published a June 30, 2026 evaluation of Claude Sonnet 5 that makes the cost story more complicated than list pricing suggests. In its methodology, Sonnet 5 scored 53 on the Intelligence Index, but under the standard-price assumptions used then it cost $2.29 per Intelligence Index task: about 2x Sonnet 4.6 and about 15% more than Opus 4.8.

The June evaluation used the planned $3/$15 rate. Anthropic retained $2/$10 as Sonnet 5’s standard global price, so those task-cost figures need that pricing context. The difference came from usage: at max effort, Artificial Analysis found Sonnet 5 used about 40% more output tokens than Sonnet 4.6 per Intelligence Index task and about 3x the agentic turns on AA-Briefcase and GDPval-AA. On GDPval-AA, max effort used about 6x more turns than low effort.

That does not make Sonnet 5 a poor choice. Artificial Analysis found it matched or outperformed Opus 4.8 on AA-Briefcase and GDPval-AA, while Opus 4.8 stayed stronger on heavier reasoning and knowledge benchmarks. The operational lesson still applies at the current standard rates: list token price and cost per completed task can point in opposite directions. Recalculate with current rates and your own measured usage.

Sonnet 5 or Opus 5 on Bedrock?

RequirementStart by testing Sonnet 5 when...Route to Opus 5 when...
High-volume production agentsCost, speed, and repeatable tool use matter more than squeezing out the last few points of reasoning accuracy.The agent makes consequential decisions where one more successful resolution is worth the premium.
Engineering workflowsThe work is reviewable: refactors, tests, documentation, triage, code search, and CI-assisted fixes.The work is ambiguous, cross-system, and expensive to correct after the fact.
Document and knowledge workThe output can be checked against source documents, citations, or structured acceptance criteria.The task requires deep judgment across conflicting evidence and weak source material.
Latency-sensitive workflowsUsers wait inside a product or operational workflow and latency is part of the user experience.Accuracy is more important than turnaround time, or the workflow is asynchronous.
FinOps postureYou need a sustainable default that can scale beyond a pilot.You have a narrow premium lane with owner approval and a measured cost-per-success advantage.

Promote Sonnet 5 to the default only if workload evaluations meet your quality, latency, and cost criteria. Route selected cases to Opus 5 when measured results justify that choice, with explicit rules for escalation and human review.

How to calculate cost per accepted result

AWS Standard prices vary by routing. The table shows the price per million input/output tokens for each model.

RoutingSonnet 5Opus 5
Global$2 / $10$5 / $25
Commercial regional$2.20 / $11$5.50 / $27.50
GovCloud$2.40 / $12$6 / $30

AWS’s model cards list Priority, Flex, and Reserved as unsupported for both models. Sonnet 5 Batch pricing is N/A; Opus 5 supports Batch. A generic service-tier setting does not establish model support.

The useful calculation is not just input and output tokens. Measure cost per accepted result: tokens, latency, effort setting, agent turns, retries, review time, tool-call failures, Opus escalations, schema rejections, and human intervention. A cheaper model can become expensive if it creates more loops. A more expensive model can be cheaper if it completes work with less review.

This is why list token price and cost per completed task can point in opposite directions. For agentic workloads, the number of turns and amount of generated output can matter as much as the published input/output token price.

Sonnet 5’s tokenizer uses approximately 30% more tokens than Sonnet 4.6 for equivalent text. Measure actual token usage at the chosen effort level before comparing costs.

What changes for AWS teams

1. Regional routing needs explicit approval

For Sonnet 5 and Opus 5, Canada Central and Calgary are Runtime geo/global source Regions, not in-Region inference locations; São Paulo supports global Runtime routing. The US geography includes US and Canadian destinations and is not Canada-only residency. Mantle lists in-Region access in us-east-1, us-gov-west-1, eu-north-1, eu-west-1, and ap-southeast-4. Verify the endpoint-specific table and approved destinations.

2. Endpoint choice affects controls

Runtime supports Invoke, Converse, and Messages for these models; Mantle supports Messages. Claude Code uses Invoke or its separate Mantle integration, not Converse. Runtime supports Bedrock Guardrails and invocation logging; Mantle does not. Choose the endpoint by the controls your workload requires.

3. Model IDs should be allowlisted

For production, avoid broad foundation-model/* or inference-profile/* permissions as the final posture. Allowlist the approved Sonnet 5 and escalation model IDs, include the routed backing models required by your inference profile, and apply Region restrictions with approved-profile exceptions for cross-Region inference. Keep Marketplace subscription, model access, and logging configuration out of day-to-day engineer runtime roles.

4. Update request parameters before migrating

Sonnet 5 enables adaptive thinking by default. Requests using manual thinking.type="enabled"/budget_tokens or non-default sampling parameters return 400 errors. Update those request fields and test your client or gateway before switching production traffic. See the Sonnet 5 migration details.

A practical Sonnet 5 rollout on Bedrock

  1. Pick one workload. Choose a real workflow with enough volume to expose cost and enough reviewability to catch failures.
  2. Freeze the baseline. Capture current model, prompt, success rate, token use, latency, retries, review time, and failure classes.
  3. Run Sonnet 5 against the same eval set. Use the exact model ID, region, endpoint, and effort setting you would use in production.
  4. Compare against Opus 5 only where the decision matters. Do not run every request through Opus just because it scores higher publicly.
  5. Set routing rules. Decide what stays on Sonnet 5, what escalates to Opus, what requires human approval, and what should not run through an LLM yet.
  6. Lock the AWS controls. IAM allowlists, region restrictions, Bedrock invocation logging where supported, budgets, anomaly alerts, and owner review should be in place before expansion.
  7. Re-price with current rates. Re-run the cost model using the current standard rates, measured usage, and approved routing before expanding production.

Where Sonnet 5 fits in agentic architecture

Sonnet 5 is not just a chat model upgrade. It changes the practical cost envelope for governed AI agent sandboxes on AWS, coding assistants on Bedrock, and internal automation agents that were too expensive or too brittle on prior defaults.

That does not remove the need for architecture. Agents still need scoped tools, constrained credentials, bounded memory, audit trails, fallback behavior, and cost controls. The model got stronger; the operating model still decides whether the workflow is safe to scale.

If the agent works inside Slack or shared channels, read the Claude Tag control-boundary guide as well: Claude Tag in Slack: how it works, what it can access, and a safe AWS rollout. The same principle applies: understand the identity, data, runtime, and cost boundaries before expanding access.

FAQ

Is Sonnet 5 cheaper than Opus 5?

Its Standard per-token prices are lower. Compare cost per accepted result as well: retries, agent turns, cache usage, tools, escalation, and review time can change the outcome.

Does Sonnet 5 replace Opus for coding?

Test both on your workload. Promote Sonnet 5 where quality, latency, and cost meet acceptance criteria; retain Opus 5 for cases where measured results justify it.

How is Sonnet 5 priced?

Standard global inference costs $2 per million input tokens and $10 per million output tokens. Supported commercial regional routing costs $2.20/$11, and GovCloud costs $2.40/$12. Include cache usage and tool costs when estimating a complete task.

Is Sonnet 5 available in Canada or Brazil?

Yes, through the supported Runtime cross-Region paths above. Source availability is not local inference or a country-only residency guarantee.

How long will these model versions remain available?

AWS lists both as Active: Sonnet 5 launched June 30, 2026, with EOL no sooner than June 30, 2027; Opus 5 launched July 24, 2026, with EOL no sooner than July 24, 2027. These are minimum lifecycle commitments, not scheduled shutdown dates.

Does 1M context need a special setting?

Sonnet 5 always uses 1M context in Claude Code and has no [1m] variant. For extended Opus context on Invoke, select the [1m] variant. Both AWS model cards list a 128K maximum output.

How Elevata can help

Bring one real workload, the current model path, and your AWS constraints. Elevata can help you benchmark Sonnet 5 against your existing workflow, compare it with Opus 5 where needed, model current costs, and review the Bedrock controls required for production.

Useful next reads: Amazon Bedrock consulting, AWS cost optimization, and the Opus 4.8 benchmark guide.

Talk to Elevata about a Sonnet 5 workload review.

Related

Continue reading

Related reading on this topic.