Skip to content

14 models priced

LLM API pricing

Grouped by what they actually cost to run rather than by vendor marketing tier. All figures are US dollars per million tokens at standard, non-batch, non-cached rates.

Pricing checked against provider documentation on . How we verify

Budget

under $1 blended 4 models

Priced for workloads measured in millions of calls. The 2026 price cuts pushed genuinely capable models into this bracket — several now outperform the previous generation's flagships.

Mid-range

$1–$5 blended 6 models

Where most production traffic belongs. These handle customer support, RAG pipelines, and mid-complexity coding without the flagship premium.

Premium

$5+ blended 4 models

Worth it when errors are expensive or a task runs for hours unsupervised. Benchmark against the mid-range first — the capability gap is often smaller than the price gap.

Full pricing table

14 of 14 models

Model Provider Input / 1M Output / 1M Context Best for Status
DeepSeek V4-Flash reasoning open DeepSeek $0.22 $0.66 1M The cheapest capable open-weight option stable
GPT-5.6 Luna reasoning OpenAI $0.20 $1.20 1.05M High-volume work on a tight budget stable
Gemini 3.1 Flash-Lite Google $0.25 $1.50 1.05M Ultra-cheap high-volume tasks stable
DeepSeek V4-Pro reasoning open DeepSeek $0.66 $1.98 1M Frontier-adjacent reasoning at open-weight prices stable
Claude Haiku 4.5 reasoning Anthropic $1.00 $5.00 200K Lowest-latency Claude responses stable
GLM-5.2 reasoning open Z.ai $1.40 $4.40 1M MIT-licensed coding agents stable
Gemini 3.5 Flash reasoning Google $1.50 $9.00 1.05M Agentic loops and sub-agent fleets stable
Claude Sonnet 5 reasoning Anthropic $2.00 $10.00 1M Fast Claude-quality reasoning at scale stable
GPT-5.6 Terra reasoning OpenAI $2.00 $12.00 1.05M Balanced production workloads stable
Gemini 3.1 Pro reasoning Google $2.00 $12.00 1.05M Best price-to-reasoning ratio preview
Kimi K3 reasoning open Moonshot AI $3.00 $15.00 1.05M Natively multimodal long-horizon agents stable
GPT-5.6 Sol reasoning OpenAI $4.00 $20.00 1.05M Frontier coding and long-horizon agents stable
Claude Opus 5 reasoning Anthropic $5.00 $25.00 1M Most teams' best Claude starting point stable
Claude Fable 5 reasoning Anthropic $10.00 $50.00 1M Long-horizon agentic engineering stable

Conditions attached to headline prices

30 of the models we track carry a condition that changes what you actually pay. These matter more than a few cents per million tokens.

  • GPT-5.6 Sol — Promotional pricing (down from $5/$30) held at least through 2026-11-21. Prompts above 272K input tokens bill at higher rates; 'fast' processing is 2x.
  • GPT-5.6 Terra — Cut 20% from the $2.50/$15 launch price on 2026-07-30.
  • GPT-5.6 Luna — Cut 80% from the $1/$6 launch price on 2026-07-30.
  • Claude Fable 5 — 5-minute cache writes $12.50/MTok, 1-hour $20/MTok. Requires 30-day data retention, so not available under zero-data-retention terms.
  • Claude Mythos 5 — Same price as Fable 5. Access restricted to Project Glasswing partners.
  • Gemini 3.1 Pro — Prompts over 200K tokens reprice to $4 input / $18 output per 1M.
  • Gemini 3.5 Flash — Flat rate at any context length, unlike Gemini 3.1 Pro's 200K cliff.
  • DeepSeek V4-Pro — Off-peak rate shown. Peak rates are exactly double ($1.32 input / $3.96 output) during 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Cache hits cost roughly 3% of a cache miss.
  • DeepSeek V4-Flash — Off-peak rate shown. Peak rates double ($0.44 / $1.32) during 01:00–04:00 and 06:00–10:00 UTC on weekdays. Superseded the earlier flat-rate pricing of $0.14 / $0.28 that many comparison sites still quote.
  • GLM-5.2 — Independently measured by Artificial Analysis; the vendor's own pricing page was unreachable at time of checking. Treat as approximate and confirm in the console.
  • Kimi K3 — The custom licence requires a separate commercial agreement once the licensee and its affiliates exceed $20M revenue over any consecutive 12 months.
  • Universal-2 — Batch transcription. Speech-understanding add-ons are billed separately.
  • GPT-4o Mini Transcribe — The OpenAI transcription endpoint caps uploads at 25 MB, which a single long recording will exceed.
  • Scribe v2 — Published as $0.22 per audio hour. Entity detection adds $0.07/hour and keyterm prompting $0.05/hour. A realtime variant runs $0.39/hour (~$0.0065/min) at roughly 150 ms latency.
  • Nova-3 — Batch pay-as-you-go rate for the monolingual model. Streaming is higher and quoted between $0.0048 and $0.0077/min depending on plan and source; multilingual adds roughly 20%. Confirm your tier in the console.
  • GPT-4o Transcribe — Legacy whisper-1 costs the same. Live transcript streaming through the realtime endpoint is far pricier at roughly $0.017/min. Uploads are capped at 25 MB.
  • Flash / Turbo TTS — Published as $0.05 per 1,000 characters. Text-to-speech bills per character, not per audio minute.
  • Multilingual v3 TTS — Published as $0.10 per 1,000 characters. Voice changer and isolator are $0.12/min, sound effects $0.12/min, dubbing $0.33–$0.50 per source minute.
  • Imagen 4 Fast — 1K resolution.
  • FLUX.2 [pro] — 1024² generation. The [pro] tier is API-only; FLUX's smaller schnell and dev variants are the openly downloadable ones.
  • GPT Image 1.5 — Medium quality at 1024px. High quality is reported between $0.13 and $0.17 per image and sources disagree on the exact figure — confirm in the console before budgeting. A cheaper Mini variant runs about $0.036 at high quality.
  • Nano Banana 2 — 1024² render. Priced at $60 per million output tokens, so cost tracks resolution: $0.045 at 512px, $0.067 at 1K, $0.101 at 2K, $0.151 at 4K. Batch mode halves all of these.
  • Nano Banana Pro — 1K–2K resolution on the standard tier; 4K is $0.24. Batch and Flex lanes cut both roughly in half, to $0.067 and $0.12.
  • Veo 3.1 Lite — 720p, video only, no audio. Quoted between $0.03 and $0.05/sec depending on route. Clips are 4, 6, or 8 seconds; access is quota-limited and regional.
  • Runway Gen-4 Turbo — Billed as 5 credits per second at $0.01 per credit. No native audio on the Gen-4 family.
  • Kling 3.0 — 720p without native audio. Adding audio takes it to roughly $0.11–$0.13/sec, and 1080p to $0.14–$0.168/sec. Clips run 3–15 seconds.
  • Runway Gen-4.5 — Billed as 12 credits per second at $0.01 per credit. No native audio on the Gen-4 family.
  • Veo 3.1 Fast — Audio included. Video-only routes are cheaper at roughly $0.08–$0.10/sec, and third-party resellers vary. Clips are 4, 6, or 8 seconds.
  • Veo 3.1 — 720p/1080p with audio included. Video-only routes run about $0.20/sec and 4K is priced separately. Clips are 4, 6, or 8 seconds; maximum 4 outputs per request.
  • Sora 2 — The API shuts down on 24 September 2026 — do not start new work on it. 720p standard; the Pro tier ran $0.30/sec at 720p and $0.70/sec at 1080p. Batch was half price. Veo 3.1 Fast and Kling 3.0 are the closest replacements.

Estimate volume, not unit price

Multiply your monthly requests by average input and output length, then compare monthly totals. A 10x difference in unit price can be irrelevant at low volume and decisive at high volume.

Watch the long-context cliffs

Some providers reprice the entire request once a prompt crosses a threshold — not just the excess. A prompt drifting from 195K to 205K tokens can double in cost with no visible change.

Reasoning breaks naive estimates

Hidden thinking tokens bill as output. A cheaper model that thinks five times longer costs more. Details here

Frequently asked

Why is output more expensive than input?

Generating tokens requires a forward pass per token, while input tokens are processed in parallel. That asymmetry is why output typically costs five to six times input — and why trimming responses saves more than trimming prompts.

What does 'blended cost' mean on this page?

A single number combining input and output price assuming a 3:1 input-to-output ratio, which is representative of chat and agent workloads. It exists only to rank models sensibly; use the calculator for your actual ratio.

Are these prices going to keep falling?

The 2026 trend was sharply downward — one budget tier fell 80% in a day, and a flagship dropped over 20% two months later. Some of those cuts are explicitly promotional with expiry dates, so architectural savings like caching and routing are more durable than betting on the next price cut.

How much can caching and batching actually save?

Cached input typically bills at about a tenth of the base rate, and batch processing is consistently 50% off across major providers. Together they often matter more than which model you pick.