Claude Sonnet 5
The workhorse Claude tier: a full million tokens of context and adaptive thinking at $2/$10, which undercuts the previous Sonnet 4.6 generation by a third on input while raising the context ceiling fivefold. The default pick for latency-sensitive products that still want Claude's writing and tool-use behaviour.
Pricing checked against provider documentation on . How we verify
claude-sonnet-5 reasoning Input
$2.00
/ 1M tokens
Output
$10.00
/ 1M tokens
Context
1M
1,000,000 tokens
Max output
128K
tokens per request
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $40.00/month on Claude Sonnet 5. Routing that same volume through the batch API brings it to $20.00. The next cheaper option, Gemini 3.5 Flash , would cost $33.00.
Where Claude Sonnet 5 fits
- Production chat and copilots
- Long-document analysis
- Tool-heavy agents needing fast turns
- Code review and PR automation
Specifications
- API model ID
- claude-sonnet-5
- Provider
- Anthropic
- Model family
- Claude 5
- Released
- April 2026
- Knowledge cutoff
- January 2026
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Native reasoning
- Yes
- Relative latency
- low
- Input modalities
- text, image
- Output modalities
- text
- Open weights
- No
- Cached input
- $0.20 / 1M tokens
- Batch discount
- 50%
Frequently asked
How much does Claude Sonnet 5 cost?
Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens, with cached input at $0.20. Batch processing is 50% cheaper.
What is Claude Sonnet 5's context window?
Claude Sonnet 5 accepts up to 1,000,000 tokens of context and can generate up to 128,000 output tokens per request.
Is Claude Sonnet 5 a reasoning model?
Yes. Claude Sonnet 5 performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.
What is Claude Sonnet 5 best for?
Fast Claude-quality reasoning at scale. The workhorse Claude tier: a full million tokens of context and adaptive thinking at $2/$10, which undercuts the previous Sonnet 4.6 generation by a third on input while raising the context ceiling fivefold. The default pick for latency-sensitive products that still want Claude's writing and tool-use behaviour.