Pricing verified September 18, 2026

AI model comparisons
with sources you can check

Compare current API prices and context limits, estimate workload cost, and understand the assumptions behind every number. No invented monthly bills or unlabeled benchmark claims.

A smaller, auditable comparison is more useful than a long unverified list

The catalog below includes only entries we could verify against provider documentation on the stated date. Product availability and prices can change, so each row links to its source.

Official sources first

Prices and model limits come from provider documentation, not affiliate roundups.

Dates and caveats

Introductory prices, cache assumptions, and vendor-published benchmarks are labeled.

Corrections welcome

Found a stale number? Send the source through our contact page and we will review it.

Read the full editorial and data methodology, or see who maintains the site on the editorial desk page.

Current API model comparison — September 2026

Prices are USD per one million input/output tokens. Context limits and pricing were checked on September 18, 2026.

Model / providerContextPublished positioningInput / outputEvidence
Claude Fable 5.1
Anthropic
1M Top-end Claude model for long-horizon autonomous agents; 1M context, 128K max output, adaptive thinking always on $10.00 / $50.00
Released September 1, 2026. Cache reads fell from $1.00 to $0.25 per 1M versus Fable 5, which Anthropic estimates cuts typical workload cost about 25% and agentic workload cost up to 45%. Batch API halves rates to $5 / $25.
Official source
Verified Sep 18
Claude Opus 5
Anthropic
1M Complex agentic coding and enterprise work; adaptive thinking and 128K maximum output $5.00 / $25.00 Official source
Verified Sep 18
Claude Sonnet 5
Anthropic
1M Fast agentic coding with a 1M context window; the default choice for most production work $2.00 / $10.00
The $2/$10 launch rate was scheduled to rise to $3/$15 on September 1, 2026. Anthropic has since stated that increase will not occur; the pricing page now lists $2/$10 as the standard rate with no expiry.
Official source
Standard rate
Claude Haiku 4.5
Anthropic
200K Fast current Claude model for high-volume and latency-sensitive work $1.00 / $5.00 Official source
Verified Sep 18
GPT-6 Astra
OpenAI
1.05M Current OpenAI flagship for computer use, coding, and multi-step professional work; 128K max output $10.00 / $50.00
Released September 3, 2026. First OpenAI model rated Critical for cybersecurity capability under its Preparedness Framework. Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Batch and Flex are half price; Fast mode is double.
Official source
New Sep 3
GPT-5.6 Sol
OpenAI
1.05M Frontier model for complex professional work, coding, research, and tool use $4.00 / $20.00
Promotional rate guaranteed at least through November 21, 2026. Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate.
Official source
Promo to Nov 21
GPT-5.6 Terra
OpenAI
1.05M Balanced GPT-5.6 tier for strong capability at a lower unit price $2.00 / $12.00
Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate.
Official source
Verified Sep 18
GPT-5.6 Luna
OpenAI
1.05M Lowest-cost GPT-5.6 tier; cheapest current-generation frontier rate published by OpenAI $0.20 / $1.20
Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate.
Official source
Verified Sep 18
Gemini 3.1 Pro
Google DeepMind
1M Google flagship for reasoning, agentic work, and long context; multimodal input $2.00 / $12.00
Prompts above 200K input tokens bill at $4.00 input and $18.00 output per 1M for the whole request. Paid tier only since April 1, 2026; the free tier no longer covers Pro models.
Official source
Verified Sep 18
Gemini 3.8 Flash
Google DeepMind
1M Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output $0.75 / $3.75
Launched September 2, 2026 on the same introductory rate as Gemini 3.6 and 3.7 Flash. Standard pricing of $1.50 / $7.50 per 1M applies from January 1, 2027. Thinking tokens bill at the output rate.
Official source
Promo to Dec 31
DeepSeek V4.1 Flash
DeepSeek
1M Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support $0.15 / $0.60
Off-peak rate shown. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $0.30 and output to $1.20 per 1M. Model name is deepseek-flash; the retired deepseek-v4-flash identifier still resolves here.
Official source
Off-peak rate
Model / providerContextPublished positioningInput / outputEvidence
DeepSeek V4.1 Flash
DeepSeek
1M Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support $0.15 / $0.60
Off-peak rate shown. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $0.30 and output to $1.20 per 1M. Model name is deepseek-flash; the retired deepseek-v4-flash identifier still resolves here.
Official source
Off-peak rate
DeepSeek V4 Pro
DeepSeek
1M Higher-capability DeepSeek model with thinking mode and 384K maximum output $0.66 / $1.98
Off-peak rate shown; peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $1.32 and output to $3.96 per 1M. DeepSeek had planned to withdraw this model on September 14, 2026 and has since extended it.
Official source
Off-peak rate
Important: token prices are not monthly subscription prices. Your bill depends on token volume, output ratio, caching, provider tiers, and any regional or platform fees. Use the cost calculator for a transparent estimate.

Four useful reference points

These cards show the exact published input/output unit rates. They do not pretend that every developer has the same “monthly cost.”

GPT-5.6 Luna
OpenAI API
$0.20 / $1.20 / 1M tokens (input/output)
Lowest published unit price in this verified set
  • 1.05M context window
  • Cached input: $0.02 / 1M tokens
  • Prompts above 272K bill at a premium
  • Official API model page linked below
Official model page ↗
Gemini 3.8 Flash
Google AI / Vertex AI
$0.75 / $3.75 / 1M tokens (input/output)
Introductory rate through December 31, 2026
  • 1M context window
  • Multimodal input
  • Vendor-published benchmark data
  • Reverts to $1.50 / $7.50 on January 1
Official model page ↗

Turn published rates into your own decision

AI API Cost Calculator

Enter your actual calls, input/output sizes, cache rate, and working days. Every assumption stays visible.

Calculate workload cost →

Token Counter & Cost Estimator

Estimate tokens for pasted text and compare the same workload across the verified catalog.

Estimate tokens →

Benchmark Comparison

Sort a deliberately small set of comparable, vendor-published coding results with the source and caveat shown.

Open benchmark matrix →

Model Shortlist Wizard

Build a transparent shortlist by workload, unit-price priority, and context need.

Build a shortlist →

Agent Blueprint Builder

Generate a reviewable blueprint covering permissions, approvals, budgets, and evaluation.

Build a blueprint →

What changed in the current catalog

New · Sep 19, 2026

The September cache price war

Five vendors repriced prompt caching within three weeks. Two flagships with identical $10/$50 list prices now bill 1.8x apart on the same agent task — the difference is entirely in the cache row. Read the worked bill →

Rechecked Sep 18, 2026

OpenAI GPT-6 Astra and the corrected GPT-5.6 rates

Astra joined on September 3 at $10/$50. Re-reading the model pages also corrected Sol to $4/$20, Terra to $2/$12, and Luna to $0.20/$1.20 — the previous values came from the launch post, and Luna differed by five times. OpenAI source ↗

Rechecked Sep 18, 2026

Claude Sonnet 5 price status

Anthropic cancelled the September 1 increase. Its pricing page now states that the $2/$10 rate "is now the standard price" and that the scheduled move to $3/$15 "will not occur". The row is no longer flagged. Anthropic source ↗

Rechecked Sep 18, 2026

DeepSeek V4.1 Flash and V4 Pro

Rates corrected to the off-peak card: Flash is $0.15/$0.60 and Pro $0.66/$1.98, each doubling during peak hours. DeepSeek also extended V4 Pro past its planned September 14 withdrawal. DeepSeek source ↗

Rechecked Sep 18, 2026

Gemini 3.8 Flash

Released September 2 on the $0.75/$3.75 introductory rate through December 31. Google now publishes a cache-read rate, so the catalog no longer assumes no discount. Google source ↗

The arguments everyone is having, settled with sources

Every claim dated, every number labeled as fact, estimate, or judgment — including the places where credible sources disagree.

Should AI Development Slow Down?

What “Pacing the Frontier” actually proposes, who supports and rejects it, and what the proposed safeguard costs against a single training run.

Read the breakdown →

Is AI Actually Dangerous?

1,663 recorded incidents sorted by cause, system type, and severity. Misuse outranks malfunction; software harms outnumber robots.

See the incident data →

What a Frontier Model Costs to Train

Estimated training costs from GPT-3 to today, where the money goes, and why chip sellers and model labs want opposite things from a slowdown.

Follow the money →

Is AI Taking Entry-Level Jobs?

Payroll records and employer surveys tell opposite stories. What each one measures, and where causation is still contested.

Compare the sources →

Practical concepts that outlast a release cycle

Model Pricing Pages

One rate card per current model with worked workload costs, plus 8 head-to-head comparisons that price the same 200-step task twice.

Browse the reference →

When Agents Attack

The PaperCut agent swarm and Google ADK's CVSS 10.0 — two verified September incidents, and what they change about agent defenses.

Read the incident record →

Choose a Model & Agent Stack

Evaluate the complete system against task outcomes, safety gates, latency, and cost per accepted result.

Read the guide →

LLM API Cost Planning

Budget for token mix, caching, retries, agent loops, tools, and real production traffic.

Read the guide →

Prompt Caching

Learn when repeated prefixes save money and how to measure misses, retention, and privacy.

Read the guide →

Local LLM Deployment

Compare Ollama, llama.cpp, and vLLM by workload, hardware, operations, and security.

Read the guide →

Fine-tuning vs RAG

A decision framework for choosing retrieval, training, or a hybrid based on the problem you actually have.

Read the guide →

Model Context Protocol

Understand MCP components, trust boundaries, and when a tool protocol helps an agent system.

Read the guide →

Structured Outputs

Use schemas and validation to make model output safer for downstream software.

Read the guide →

AI Agent Security

Threat-model tool access, secrets, untrusted context, and human approval boundaries.

Read the guide →

What the numbers do—and do not—mean

No. They are provider-published rates per one million tokens. Monthly cost depends on your workload.
No. The benchmark matrix clearly labels vendor-published results and links the source. We do not present them as independent lab results.
Each data snapshot shows its verification date. We recheck primary sources during updates and accept documented correction reports.

Model prices change without notice

Rates in this category moved repeatedly over the past year. The fastest way to catch a change is the feed, which carries every new and updated guide.

RSS feed

Subscribe in any reader. No email address, no tracking, no third party.

Open feed.xml →

Email digest

Set a reminder to re-check the catalog, or send us a note and we will tell you when the pricing review is refreshed.

Request an update →

Corrections

Every figure carries a check date and a source link. If one is wrong, the fastest fix is a message with the source.

Report a figure →