Skip to content

Claude Haiku 4.5

The fastest model Anthropic ships, still on the Claude 4.5 generation with a 200K window and an early-2025 knowledge cutoff. Worth it when response latency is the product requirement; otherwise Sonnet 5 offers more capability and five times the context for twice the price.

Pricing checked against provider documentation on . How we verify

stable Anthropic Claude 4.5 claude-haiku-4-5 reasoning

Input

$1.00

/ 1M tokens

Output

$5.00

/ 1M tokens

Context

200K

200,000 tokens

Max output

64K

tokens per request

What it costs in practice

A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $20.00/month on Claude Haiku 4.5. Routing that same volume through the batch API brings it to $10.00. The next cheaper option, DeepSeek V4-Pro , would cost $10.56.

Run your own numbers in the calculator →

Where Claude Haiku 4.5 fits

  • Real-time chat suggestions
  • Lightweight extraction and tagging
  • Guardrail and routing calls
  • High-frequency background jobs

Specifications

API model ID
claude-haiku-4-5
Provider
Anthropic
Model family
Claude 4.5
Released
October 2025
Knowledge cutoff
February 2025
Context window
200,000 tokens
Max output
64,000 tokens
Native reasoning
Yes
Relative latency
low
Input modalities
text, image
Output modalities
text
Open weights
No
Cached input
$0.10 / 1M tokens
Batch discount
50%

Frequently asked

How much does Claude Haiku 4.5 cost?

Claude Haiku 4.5 costs $1.00 per million input tokens and $5.00 per million output tokens, with cached input at $0.10. Batch processing is 50% cheaper.

What is Claude Haiku 4.5's context window?

Claude Haiku 4.5 accepts up to 200,000 tokens of context and can generate up to 64,000 output tokens per request.

Is Claude Haiku 4.5 a reasoning model?

Yes. Claude Haiku 4.5 performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.

What is Claude Haiku 4.5 best for?

Lowest-latency Claude responses. The fastest model Anthropic ships, still on the Claude 4.5 generation with a 200K window and an early-2025 knowledge cutoff. Worth it when response latency is the product requirement; otherwise Sonnet 5 offers more capability and five times the context for twice the price.