Verified August 5, 2026

AI model comparisons
with sources you can check

Compare current API prices and context limits, estimate workload cost, and understand the assumptions behind every number. No invented monthly bills or unlabeled benchmark claims.

A smaller, auditable comparison is more useful than a long unverified list

The catalog below includes only entries we could verify against provider documentation on the stated date. Product availability and prices can change, so each row links to its source.

Official sources first

Prices and model limits come from provider documentation, not affiliate roundups.

Dates and caveats

Introductory prices, cache assumptions, and vendor-published benchmarks are labeled.

Corrections welcome

Found a stale number? Send the source through our contact page and we will review it.

Read the full editorial and data methodology, or see who maintains the site on the editorial desk page.

Current API model comparison — August 2026

Prices are USD per one million input/output tokens. Context limits and pricing were checked on August 5, 2026.

Model / providerContextPublished positioningInput / outputEvidence
Claude Opus 5
Anthropic
1M Complex agentic coding and enterprise work; adaptive thinking and 128K maximum output $5.00 / $25.00 Official source
Verified Aug 5
Claude Sonnet 5
Anthropic
1M Fast agentic coding with a 1M context window; introductory price through Aug 31, 2026 $2.00 / $10.00
Standard price becomes $3 input / $15 output after August 31, 2026.
Official source
Intro price
GPT-5.6 Sol
OpenAI
1.05M Frontier model for complex professional work, coding, research, and tool use $5.00 / $30.00 Official source
Verified Aug 5
GPT-5.6 Terra
OpenAI
1.05M Balanced GPT-5.6 tier for strong capability at a lower unit price $2.50 / $15.00 Official source
Verified Aug 5
GPT-5.6 Luna
OpenAI
1.05M Cost-sensitive, high-volume GPT-5.6 tier with the same published context limit $1.00 / $6.00 Official source
Verified Aug 5
Gemini 3.6 Flash
Google DeepMind
1M Multimodal model for coding, knowledge work, and long-context understanding $1.50 / $7.50 Official source
Verified Aug 5
DeepSeek V4 Flash
DeepSeek
1M Low-cost current DeepSeek API model with thinking mode and Responses API support $0.14 / $0.28
Peak/off-peak pricing was announced without an effective date when checked.
Official source
Verified Aug 5
Claude Haiku 4.5
Anthropic
200K Fast current Claude model for high-volume and latency-sensitive work $1.00 / $5.00 Official source
Verified Aug 5
Model / providerContextPublished positioningInput / outputEvidence
DeepSeek V4 Flash
DeepSeek
1M Low-cost current DeepSeek API model with thinking mode and Responses API support $0.14 / $0.28
Peak/off-peak pricing was announced without an effective date when checked.
Official source
Verified Aug 5
DeepSeek V4 Pro
DeepSeek
1M Higher-capability DeepSeek V4 API model with thinking mode and 384K maximum output $0.43 / $0.87
Peak/off-peak pricing was announced without an effective date when checked.
Official source
Verified Aug 5
Important: token prices are not monthly subscription prices. Your bill depends on token volume, output ratio, caching, provider tiers, and any regional or platform fees. Use the cost calculator for a transparent estimate.

Four useful reference points

These cards show the exact published input/output unit rates. They do not pretend that every developer has the same “monthly cost.”

DeepSeek V4 Flash
DeepSeek API
$0.14 / $0.28 / 1M tokens (input/output)
Lowest published unit price in this verified set
  • 1M context window
  • Cache-hit input: $0.0028 / 1M tokens
  • Thinking mode available
  • Official API pricing linked below
Official pricing ↗
GPT-5.6 Luna
OpenAI API
$1.00 / $6.00 / 1M tokens (input/output)
Lower-cost GPT-5.6 tier for volume workloads
  • 1.05M context window
  • $0.10 cached input / 1M tokens
  • Tool use and coding support
  • Official API model page linked below
Official model page ↗
Gemini 3.6 Flash
Google AI / Vertex AI
$1.50 / $7.50 / 1M tokens (input/output)
Multimodal, long-context option
  • 1M context window
  • Multimodal input
  • Vendor-published benchmark data
  • Calculator assumes no cache discount
Official model page ↗

Turn published rates into your own decision

AI API Cost Calculator

Enter your actual calls, input/output sizes, cache rate, and working days. Every assumption stays visible.

Calculate workload cost →

Token Counter & Cost Estimator

Estimate tokens for pasted text and compare the same workload across the verified catalog.

Estimate tokens →

Benchmark Comparison

Sort a deliberately small set of comparable, vendor-published coding results with the source and caveat shown.

Open benchmark matrix →

Model Shortlist Wizard

Build a transparent shortlist by workload, unit-price priority, and context need.

Build a shortlist →

Agent Blueprint Builder

Generate a reviewable blueprint covering permissions, approvals, budgets, and evaluation.

Build a blueprint →

What changed in the current catalog

Checked Aug 5, 2026

OpenAI GPT-5.6 tiers

Sol, Terra, and Luna are listed separately so readers can compare published unit rates without collapsing them into one generic “GPT” price. OpenAI source ↗

Checked Aug 5, 2026

Claude Sonnet 5 introductory price

The catalog labels the temporary $2/$10 rate and the announced $3/$15 price after August 31. Anthropic source ↗

Checked Aug 5, 2026

DeepSeek V4 pricing

Flash and Pro are separated, including cache-hit rates and the provider's not-yet-effective peak/off-peak note. DeepSeek source ↗

Checked Aug 5, 2026

Gemini 3.6 Flash

The model card uses the provider-published 1M context figure and does not assume an undocumented cache discount. Google source ↗

Practical concepts that outlast a release cycle

Choose a Model & Agent Stack

Evaluate the complete system against task outcomes, safety gates, latency, and cost per accepted result.

Read the guide →

LLM API Cost Planning

Budget for token mix, caching, retries, agent loops, tools, and real production traffic.

Read the guide →

Prompt Caching

Learn when repeated prefixes save money and how to measure misses, retention, and privacy.

Read the guide →

Local LLM Deployment

Compare Ollama, llama.cpp, and vLLM by workload, hardware, operations, and security.

Read the guide →

Fine-tuning vs RAG

A decision framework for choosing retrieval, training, or a hybrid based on the problem you actually have.

Read the guide →

Model Context Protocol

Understand MCP components, trust boundaries, and when a tool protocol helps an agent system.

Read the guide →

Structured Outputs

Use schemas and validation to make model output safer for downstream software.

Read the guide →

AI Agent Security

Threat-model tool access, secrets, untrusted context, and human approval boundaries.

Read the guide →

What the numbers do—and do not—mean

No. They are provider-published rates per one million tokens. Monthly cost depends on your workload.
No. The benchmark matrix clearly labels vendor-published results and links the source. We do not present them as independent lab results.
Each data snapshot shows its verification date. We recheck primary sources during updates and accept documented correction reports.