Official sources first
Prices and model limits come from provider documentation, not affiliate roundups.
Compare current API prices and context limits, estimate workload cost, and understand the assumptions behind every number. No invented monthly bills or unlabeled benchmark claims.
The catalog below includes only entries we could verify against provider documentation on the stated date. Product availability and prices can change, so each row links to its source.
Prices and model limits come from provider documentation, not affiliate roundups.
Introductory prices, cache assumptions, and vendor-published benchmarks are labeled.
Found a stale number? Send the source through our contact page and we will review it.
Prices are USD per one million input/output tokens. Context limits and pricing were checked on September 18, 2026.
| Model / provider | Context | Published positioning | Input / output | Evidence |
|---|---|---|---|---|
| Claude Fable 5.1 Anthropic |
1M | Top-end Claude model for long-horizon autonomous agents; 1M context, 128K max output, adaptive thinking always on | $10.00 / $50.00 Released September 1, 2026. Cache reads fell from $1.00 to $0.25 per 1M versus Fable 5, which Anthropic estimates cuts typical workload cost about 25% and agentic workload cost up to 45%. Batch API halves rates to $5 / $25. |
Official source Verified Sep 18 |
| Claude Opus 5 Anthropic |
1M | Complex agentic coding and enterprise work; adaptive thinking and 128K maximum output | $5.00 / $25.00 | Official source Verified Sep 18 |
| Claude Sonnet 5 Anthropic |
1M | Fast agentic coding with a 1M context window; the default choice for most production work | $2.00 / $10.00 The $2/$10 launch rate was scheduled to rise to $3/$15 on September 1, 2026. Anthropic has since stated that increase will not occur; the pricing page now lists $2/$10 as the standard rate with no expiry. |
Official source Standard rate |
| Claude Haiku 4.5 Anthropic |
200K | Fast current Claude model for high-volume and latency-sensitive work | $1.00 / $5.00 | Official source Verified Sep 18 |
| GPT-6 Astra OpenAI |
1.05M | Current OpenAI flagship for computer use, coding, and multi-step professional work; 128K max output | $10.00 / $50.00 Released September 3, 2026. First OpenAI model rated Critical for cybersecurity capability under its Preparedness Framework. Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Batch and Flex are half price; Fast mode is double. |
Official source New Sep 3 |
| GPT-5.6 Sol OpenAI |
1.05M | Frontier model for complex professional work, coding, research, and tool use | $4.00 / $20.00 Promotional rate guaranteed at least through November 21, 2026. Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate. |
Official source Promo to Nov 21 |
| GPT-5.6 Terra OpenAI |
1.05M | Balanced GPT-5.6 tier for strong capability at a lower unit price | $2.00 / $12.00 Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate. |
Official source Verified Sep 18 |
| GPT-5.6 Luna OpenAI |
1.05M | Lowest-cost GPT-5.6 tier; cheapest current-generation frontier rate published by OpenAI | $0.20 / $1.20 Prompts over 272K input tokens bill at 2x input and 1.5x output for the whole request. Cache writes bill at 1.25x the uncached input rate. |
Official source Verified Sep 18 |
| Gemini 3.1 Pro Google DeepMind |
1M | Google flagship for reasoning, agentic work, and long context; multimodal input | $2.00 / $12.00 Prompts above 200K input tokens bill at $4.00 input and $18.00 output per 1M for the whole request. Paid tier only since April 1, 2026; the free tier no longer covers Pro models. |
Official source Verified Sep 18 |
| Gemini 3.8 Flash Google DeepMind |
1M | Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output | $0.75 / $3.75 Launched September 2, 2026 on the same introductory rate as Gemini 3.6 and 3.7 Flash. Standard pricing of $1.50 / $7.50 per 1M applies from January 1, 2027. Thinking tokens bill at the output rate. |
Official source Promo to Dec 31 |
| DeepSeek V4.1 Flash DeepSeek |
1M | Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support | $0.15 / $0.60 Off-peak rate shown. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $0.30 and output to $1.20 per 1M. Model name is deepseek-flash; the retired deepseek-v4-flash identifier still resolves here. |
Official source Off-peak rate |
| Model / provider | Context | Published positioning | Input / output | Evidence |
|---|---|---|---|---|
| DeepSeek V4.1 Flash DeepSeek |
1M | Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support | $0.15 / $0.60 Off-peak rate shown. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $0.30 and output to $1.20 per 1M. Model name is deepseek-flash; the retired deepseek-v4-flash identifier still resolves here. |
Official source Off-peak rate |
| DeepSeek V4 Pro DeepSeek |
1M | Higher-capability DeepSeek model with thinking mode and 384K maximum output | $0.66 / $1.98 Off-peak rate shown; peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $1.32 and output to $3.96 per 1M. DeepSeek had planned to withdraw this model on September 14, 2026 and has since extended it. |
Official source Off-peak rate |
These cards show the exact published input/output unit rates. They do not pretend that every developer has the same “monthly cost.”
Enter your actual calls, input/output sizes, cache rate, and working days. Every assumption stays visible.
Calculate workload cost →Estimate tokens for pasted text and compare the same workload across the verified catalog.
Estimate tokens →Sort a deliberately small set of comparable, vendor-published coding results with the source and caveat shown.
Open benchmark matrix →Build a transparent shortlist by workload, unit-price priority, and context need.
Build a shortlist →Generate a reviewable blueprint covering permissions, approvals, budgets, and evaluation.
Build a blueprint →Five vendors repriced prompt caching within three weeks. Two flagships with identical $10/$50 list prices now bill 1.8x apart on the same agent task — the difference is entirely in the cache row. Read the worked bill →
Astra joined on September 3 at $10/$50. Re-reading the model pages also corrected Sol to $4/$20, Terra to $2/$12, and Luna to $0.20/$1.20 — the previous values came from the launch post, and Luna differed by five times. OpenAI source ↗
Anthropic cancelled the September 1 increase. Its pricing page now states that the $2/$10 rate "is now the standard price" and that the scheduled move to $3/$15 "will not occur". The row is no longer flagged. Anthropic source ↗
Rates corrected to the off-peak card: Flash is $0.15/$0.60 and Pro $0.66/$1.98, each doubling during peak hours. DeepSeek also extended V4 Pro past its planned September 14 withdrawal. DeepSeek source ↗
Released September 2 on the $0.75/$3.75 introductory rate through December 31. Google now publishes a cache-read rate, so the catalog no longer assumes no discount. Google source ↗
Every claim dated, every number labeled as fact, estimate, or judgment — including the places where credible sources disagree.
What “Pacing the Frontier” actually proposes, who supports and rejects it, and what the proposed safeguard costs against a single training run.
Read the breakdown →1,663 recorded incidents sorted by cause, system type, and severity. Misuse outranks malfunction; software harms outnumber robots.
See the incident data →Estimated training costs from GPT-3 to today, where the money goes, and why chip sellers and model labs want opposite things from a slowdown.
Follow the money →Payroll records and employer surveys tell opposite stories. What each one measures, and where causation is still contested.
Compare the sources →One rate card per current model with worked workload costs, plus 8 head-to-head comparisons that price the same 200-step task twice.
Browse the reference →The PaperCut agent swarm and Google ADK's CVSS 10.0 — two verified September incidents, and what they change about agent defenses.
Read the incident record →Evaluate the complete system against task outcomes, safety gates, latency, and cost per accepted result.
Read the guide →Budget for token mix, caching, retries, agent loops, tools, and real production traffic.
Read the guide →Learn when repeated prefixes save money and how to measure misses, retention, and privacy.
Read the guide →Compare Ollama, llama.cpp, and vLLM by workload, hardware, operations, and security.
Read the guide →A decision framework for choosing retrieval, training, or a hybrid based on the problem you actually have.
Read the guide →Understand MCP components, trust boundaries, and when a tool protocol helps an agent system.
Read the guide →Use schemas and validation to make model output safer for downstream software.
Read the guide →Threat-model tool access, secrets, untrusted context, and human approval boundaries.
Read the guide →Rates in this category moved repeatedly over the past year. The fastest way to catch a change is the feed, which carries every new and updated guide.
Set a reminder to re-check the catalog, or send us a note and we will tell you when the pricing review is refreshed.
Every figure carries a check date and a source link. If one is wrong, the fastest fix is a message with the source.