Skip to content

Claude Sonnet 5

The workhorse Claude tier: a full million tokens of context and adaptive thinking at $2/$10, which undercuts the previous Sonnet 4.6 generation by a third on input while raising the context ceiling fivefold. The default pick for latency-sensitive products that still want Claude's writing and tool-use behaviour.

Pricing checked against provider documentation on . How we verify

stable Anthropic Claude 5 claude-sonnet-5 reasoning

Input

$2.00

/ 1M tokens

Output

$10.00

/ 1M tokens

Context

1M

1,000,000 tokens

Max output

128K

tokens per request

What it costs in practice

A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $40.00/month on Claude Sonnet 5. Routing that same volume through the batch API brings it to $20.00. The next cheaper option, Gemini 3.5 Flash , would cost $33.00.

Run your own numbers in the calculator →

Where Claude Sonnet 5 fits

  • Production chat and copilots
  • Long-document analysis
  • Tool-heavy agents needing fast turns
  • Code review and PR automation

Specifications

API model ID
claude-sonnet-5
Provider
Anthropic
Model family
Claude 5
Released
April 2026
Knowledge cutoff
January 2026
Context window
1,000,000 tokens
Max output
128,000 tokens
Native reasoning
Yes
Relative latency
low
Input modalities
text, image
Output modalities
text
Open weights
No
Cached input
$0.20 / 1M tokens
Batch discount
50%

Frequently asked

How much does Claude Sonnet 5 cost?

Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens, with cached input at $0.20. Batch processing is 50% cheaper.

What is Claude Sonnet 5's context window?

Claude Sonnet 5 accepts up to 1,000,000 tokens of context and can generate up to 128,000 output tokens per request.

Is Claude Sonnet 5 a reasoning model?

Yes. Claude Sonnet 5 performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.

What is Claude Sonnet 5 best for?

Fast Claude-quality reasoning at scale. The workhorse Claude tier: a full million tokens of context and adaptive thinking at $2/$10, which undercuts the previous Sonnet 4.6 generation by a third on input while raising the context ceiling fivefold. The default pick for latency-sensitive products that still want Claude's writing and tool-use behaviour.