Official sources first
Prices and model limits come from provider documentation, not affiliate roundups.
Compare current API prices and context limits, estimate workload cost, and understand the assumptions behind every number. No invented monthly bills or unlabeled benchmark claims.
The catalog below includes only entries we could verify against provider documentation on the stated date. Product availability and prices can change, so each row links to its source.
Prices and model limits come from provider documentation, not affiliate roundups.
Introductory prices, cache assumptions, and vendor-published benchmarks are labeled.
Found a stale number? Send the source through our contact page and we will review it.
Prices are USD per one million input/output tokens. Context limits and pricing were checked on August 5, 2026.
| Model / provider | Context | Published positioning | Input / output | Evidence |
|---|---|---|---|---|
| Claude Opus 5 Anthropic |
1M | Complex agentic coding and enterprise work; adaptive thinking and 128K maximum output | $5.00 / $25.00 | Official source Verified Aug 5 |
| Claude Sonnet 5 Anthropic |
1M | Fast agentic coding with a 1M context window; introductory price through Aug 31, 2026 | $2.00 / $10.00 Standard price becomes $3 input / $15 output after August 31, 2026. |
Official source Intro price |
| GPT-5.6 Sol OpenAI |
1.05M | Frontier model for complex professional work, coding, research, and tool use | $5.00 / $30.00 | Official source Verified Aug 5 |
| GPT-5.6 Terra OpenAI |
1.05M | Balanced GPT-5.6 tier for strong capability at a lower unit price | $2.50 / $15.00 | Official source Verified Aug 5 |
| GPT-5.6 Luna OpenAI |
1.05M | Cost-sensitive, high-volume GPT-5.6 tier with the same published context limit | $1.00 / $6.00 | Official source Verified Aug 5 |
| Gemini 3.6 Flash Google DeepMind |
1M | Multimodal model for coding, knowledge work, and long-context understanding | $1.50 / $7.50 | Official source Verified Aug 5 |
| DeepSeek V4 Flash DeepSeek |
1M | Low-cost current DeepSeek API model with thinking mode and Responses API support | $0.14 / $0.28 Peak/off-peak pricing was announced without an effective date when checked. |
Official source Verified Aug 5 |
| Claude Haiku 4.5 Anthropic |
200K | Fast current Claude model for high-volume and latency-sensitive work | $1.00 / $5.00 | Official source Verified Aug 5 |
| Model / provider | Context | Published positioning | Input / output | Evidence |
|---|---|---|---|---|
| DeepSeek V4 Flash DeepSeek |
1M | Low-cost current DeepSeek API model with thinking mode and Responses API support | $0.14 / $0.28 Peak/off-peak pricing was announced without an effective date when checked. |
Official source Verified Aug 5 |
| DeepSeek V4 Pro DeepSeek |
1M | Higher-capability DeepSeek V4 API model with thinking mode and 384K maximum output | $0.43 / $0.87 Peak/off-peak pricing was announced without an effective date when checked. |
Official source Verified Aug 5 |
These cards show the exact published input/output unit rates. They do not pretend that every developer has the same “monthly cost.”
Enter your actual calls, input/output sizes, cache rate, and working days. Every assumption stays visible.
Calculate workload cost →Estimate tokens for pasted text and compare the same workload across the verified catalog.
Estimate tokens →Sort a deliberately small set of comparable, vendor-published coding results with the source and caveat shown.
Open benchmark matrix →Build a transparent shortlist by workload, unit-price priority, and context need.
Build a shortlist →Generate a reviewable blueprint covering permissions, approvals, budgets, and evaluation.
Build a blueprint →Sol, Terra, and Luna are listed separately so readers can compare published unit rates without collapsing them into one generic “GPT” price. OpenAI source ↗
The catalog labels the temporary $2/$10 rate and the announced $3/$15 price after August 31. Anthropic source ↗
Flash and Pro are separated, including cache-hit rates and the provider's not-yet-effective peak/off-peak note. DeepSeek source ↗
The model card uses the provider-published 1M context figure and does not assume an undocumented cache discount. Google source ↗
Evaluate the complete system against task outcomes, safety gates, latency, and cost per accepted result.
Read the guide →Budget for token mix, caching, retries, agent loops, tools, and real production traffic.
Read the guide →Learn when repeated prefixes save money and how to measure misses, retention, and privacy.
Read the guide →Compare Ollama, llama.cpp, and vLLM by workload, hardware, operations, and security.
Read the guide →A decision framework for choosing retrieval, training, or a hybrid based on the problem you actually have.
Read the guide →Understand MCP components, trust boundaries, and when a tool protocol helps an agent system.
Read the guide →Use schemas and validation to make model output safer for downstream software.
Read the guide →Threat-model tool access, secrets, untrusted context, and human approval boundaries.
Read the guide →