TL;DR: Five pricing events in five weeks. OpenAI shipped GPT-6 Astra on September 3 at $10 in / $50 out, exactly matching Anthropic’s Claude Fable 5.1, which launched September 1 at the same list price with cache reads cut 75%. Google launched Gemini 3.8 Flash on September 2 at $0.75 / $3.75 until December 31, then doubles. Anthropic made Sonnet 5’s $2 / $10 permanent and cancelled the September 1 increase. And DeepSeek, the budget benchmark everyone quoted, roughly tripled its V4 prices on August 16 and added peak-hour billing. The 143x gap we reported in June is now 68x, and it closed from the bottom, not the top.
What changed since June
If you set your model budget off our June table, most of it is now wrong. For the per-vendor deep dives see the AI Pricing Watch hub, the Anthropic API pricing page and the OpenAI vs Anthropic vs Google head-to-head.
| Date | Vendor | Change | Source |
|---|---|---|---|
| Jul 30 | OpenAI | GPT-5.6 Luna cut 80% to $0.20 / $1.20. Terra cut 20% to $2 / $12. | OpenAI pricing |
| Aug 10 | Anthropic | Sonnet 5’s introductory $2 / $10 made permanent. Scheduled $3 / $15 increase on Sep 1 cancelled. | Anthropic pricing |
| Aug 16 | DeepSeek | V4 Flash from $0.14 / $0.28 to $0.22 / $0.66 off-peak and $0.44 / $1.32 peak. V4 Pro from $0.44 / $0.87 to $0.66 / $1.98 off-peak and $1.32 / $3.96 peak. | DeepSeek pricing |
| Sep 1 | Anthropic | Claude Fable 5.1 at $10 / $50, cache reads $0.25 (was $1 on Fable 5). | Anthropic pricing |
| Sep 2 | Gemini 3.8 Flash at $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50. | Gemini pricing | |
| Sep 3 | OpenAI | GPT-6 Astra at $10 / $50, cached input $1. API access rolling out over the following days. | OpenAI pricing |
Prices are per million tokens, input / output, and were read from the vendor pages on September 4, 2026.
The full table
All prices per 1M tokens. “1M+1M” is the cost of one million input plus one million output tokens, the same shorthand we used in June so you can compare directly.
| Model | Input | Cached input | Output | 1M+1M | Tier |
|---|---|---|---|---|---|
| GPT-5 nano | $0.05 | $0.005 | $0.40 | $0.45 | Budget |
| DeepSeek V4 Flash (off-peak) | $0.22 | $0.007 | $0.66 | $0.88 | Budget |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $1.40 | Budget |
| GPT-5.4 nano | $0.20 | $0.02 | $1.25 | $1.45 | Budget |
| DeepSeek V4 Flash (peak) | $0.44 | $0.014 | $1.32 | $1.76 | Budget |
| Gemini 3.1 Flash-Lite | $0.25 | n/a | $1.50 | $1.75 | Budget |
| DeepSeek V4 Pro (off-peak) | $0.66 | $0.022 | $1.98 | $2.64 | Budget |
| Gemini 3.5 Flash-Lite | $0.30 | n/a | $2.50 | $2.80 | Budget |
| Grok 4.3 | $1.25 | n/a | $2.50 | $3.75 | Budget |
| Gemini 3.8 Flash (to Dec 31) | $0.75 | n/a | $3.75 | $4.50 | Budget |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | $5.25 | Budget |
| DeepSeek V4 Pro (peak) | $1.32 | $0.044 | $3.96 | $5.28 | Budget |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | $6.00 | Mid |
| Grok 4.6 | $2.00 | n/a | $6.00 | $8.00 | Mid |
| Gemini 3.8 Flash (from Jan 1) | $1.50 | n/a | $7.50 | $9.00 | Mid |
| Gemini 3.5 Flash | $1.50 | n/a | $9.00 | $10.50 | Mid |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | $12.00 | Mid |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $14.00 | Mid |
| Gemini 3.1 Pro (under 200K) | $2.00 | n/a | $12.00 | $14.00 | Mid |
| GPT-5.4 | $2.50 | $0.25 | $15.00 | $17.50 | Premium |
| Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | $18.00 | Premium |
| GPT-5.6 Sol (promo to Nov 21) | $4.00 | $0.40 | $20.00 | $24.00 | Premium |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | $30.00 | Frontier |
| Claude Opus 4.8 | $5.00 | $0.50 | $25.00 | $30.00 | Frontier |
| GPT-5.5 | $5.00 | $0.50 | $30.00 | $35.00 | Frontier |
| Claude Fable 5 | $10.00 | $1.00 | $50.00 | $60.00 | Ultra |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | $60.00 | Ultra |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | $60.00 | Ultra |
| GPT-5.5 Pro | $30.00 | n/a | $180.00 | $210.00 | Ultra |
Notes on the table:
- Google does not publish a separate cached-input rate on the Gemini API pricing page, so that column is blank. Gemini 3.1 Pro and the Grok models charge roughly double for prompts over 200K tokens.
- DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else is off-peak at half price. That is Asian business hours, so if you run US-hours batch jobs you mostly pay the off-peak rate.
- Anthropic and OpenAI both offer 50% off for batch or flex processing. Anthropic charges 1.25x input for a five-minute cache write and 2x for a one-hour write. OpenAI cache writes run 1.25x to 1.5x.
- Grok and the Gemini 2.5 series are listed but not discussed further. They are not price leaders in any tier.
The 68x gap, and why it shrank from the bottom
In June the spread between the cheapest usable model and the most expensive was 143x: DeepSeek V4 Flash at $0.42 per 1M+1M versus Claude Fable 5 at $60.
Today it is 68x. The top did not move. Fable 5.1 and GPT-6 Astra both landed at exactly $60. The bottom moved up, because DeepSeek’s off-peak Flash is now $0.88 and its peak rate is $1.76.
That is the single most important shift in this update. The “DeepSeek is 100x cheaper” line that anchored every cost comparison for a year no longer holds. At peak hours DeepSeek V4 Flash costs more than OpenAI’s GPT-5.6 Luna, and Luna is a first-party model from a US vendor with no off-peak fine print. Third-party marketplaces are reported to resell DeepSeek V4 Flash below list (Digital Applied’s August tracker cites $0.09 / $0.18), but that is a reseller quote, not a vendor rate, and we have not verified it.
If you built anything on the assumption that DeepSeek would stay at $0.28 output forever, re-run the numbers.
What you actually pay per task
Raw token prices mislead. Below is the list-price cost of four realistic workloads. Token assumptions are stated so you can swap in your own. All figures use standard (not cached, not batch) rates, DeepSeek at off-peak, and Gemini 3.8 Flash at the introductory rate.
| Task (input / output tokens) | DeepSeek V4 Flash | GPT-5.6 Luna | Gemini 3.8 Flash | Claude Sonnet 5 | GPT-5.6 Terra | Claude Opus 5 | GPT-5.5 | Fable 5.1 or GPT-6 Astra |
|---|---|---|---|---|---|---|---|---|
| Chatbot reply (400 / 100) | $0.00015 | $0.0002 | $0.0007 | $0.0018 | $0.0020 | $0.0045 | $0.0050 | $0.0090 |
| Document analysis (10K / 1K) | $0.0029 | $0.0032 | $0.011 | $0.030 | $0.032 | $0.075 | $0.080 | $0.15 |
| Agentic coding session (80K / 20K) | $0.031 | $0.040 | $0.14 | $0.36 | $0.40 | $0.90 | $1.00 | $1.80 |
| Codebase migration (800K / 200K) | $0.31 | $0.40 | $1.35 | $3.60 | $4.00 | $9.00 | $10.00 | $18.00 |
Two things to notice:
- The budget tier has collapsed to a rounding error. A chatbot reply on Luna or DeepSeek costs two hundredths of a cent. At that level the model’s price is irrelevant and the cost of your own infrastructure dominates. Stop optimising it.
- The frontier tier is where per-task math actually matters. A codebase migration on Fable 5.1 or Astra costs $18 versus $3.60 on Sonnet 5. If the frontier model finishes in one attempt and Sonnet takes three, Sonnet still wins on price ($10.80) but loses on wall-clock time and review burden. This is exactly the calculation the ultra tier is priced against, and it only pays off on tasks where a failed attempt costs more than $15 of your time.
The cache-read story nobody is pricing in
The headline rates for Fable 5.1 and GPT-6 Astra are identical. The cache rates are not.
| Model | Input | Cached input | Cache discount |
|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | 97.5% |
| GPT-6 Astra | $10.00 | $1.00 | 90% |
| Claude Fable 5 | $10.00 | $1.00 | 90% |
| Claude Opus 5 | $5.00 | $0.50 | 90% |
| GPT-5.5 | $5.00 | $0.50 | 90% |
For a one-shot prompt this changes nothing. For an agent that re-reads a 100K-token context on every turn it changes everything. Fifty turns over the same 100K context:
- Fable 5 or GPT-6 Astra: 5M cached tokens at $1 = $5.00 in cached input
- Fable 5.1: 5M cached tokens at $0.25 = $1.25
Anthropic’s launch material claims typical agentic workloads run 25% to 45% cheaper on 5.1 than on 5 at the same list price, and the cache line is where that comes from. Whether OpenAI matches it is the thing to watch in Astra’s first pricing revision. Until then, for cache-heavy agent loops the two $60 models are not the same price. Fable 5.1 is materially cheaper.
The three tiers that matter now
Budget: under $6 per 1M+1M
Models: GPT-5.6 Luna, DeepSeek V4 Flash, Gemini 3.8 Flash (until December 31), GPT-5.4 mini, Claude Haiku 4.5.
Best for: classification, extraction, routing, first-line support, anything with a clear pattern and a high volume.
The play: Luna is the new default. At $0.20 / $1.20 it is cheaper than DeepSeek at peak, comes from a first-party US vendor, and has no off-peak calendar to manage. Gemini 3.8 Flash at $0.75 / $3.75 is the capability upgrade inside the tier, but it doubles on January 1, so do not build a 2027 budget on the introductory rate.
Mid: $8 to $18 per 1M+1M
Models: Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.4, Claude Sonnet 4.6.
Best for: production workloads with real reasoning: document analysis, code review, RAG, content generation.
The play: Sonnet 5 at $2 / $10 is now the permanent mid-tier anchor, and it is $2 cheaper on output than Terra for the same input price. Sonnet 4.6 at $3 / $15 is the same family at a 50% premium. Unless you have a compatibility reason, there is no case for 4.6 over 5 anymore.
Frontier and ultra: $30 to $60 per 1M+1M
Models: Claude Opus 5, Claude Opus 4.8, GPT-5.5, Claude Fable 5.1, GPT-6 Astra.
Best for: multi-hour autonomous coding, legal and financial analysis, anything where one wrong answer costs real money.
The play: Opus 5 and GPT-5.5 at $30 and $35 do 90% of what the $60 models do. Pay for Fable 5.1 or Astra only on tasks where you have measured that the cheaper model fails, and if the task is cache-heavy, Fable 5.1’s $0.25 cache read makes it the cheaper of the two ultra models by a wide margin.
The smart routing math, updated
Nobody running production AI should run one model. Here is the September version of the routing example from June, using a team pushing 10M tokens a day (8M input, 2M output) with an 80 / 15 / 5 split by difficulty.
| Route | Share | Model | Daily cost |
|---|---|---|---|
| Simple | 80% | GPT-5.6 Luna | $3.20 |
| Standard | 15% | Claude Sonnet 5 | $5.40 |
| Hard | 5% | Claude Opus 5 | $4.50 |
| Routed total | $13.10 per day, about $393 per month | ||
| Everything on Sonnet 5 | 100% | $36 per day, about $1,080 per month | |
| Everything on Fable 5.1 or Astra | 100% | $180 per day, about $5,400 per month |
Routing saves 93% against running the ultra model on everything and 64% against running Sonnet 5 on everything. Caching, batch and off-peak pricing would cut all three rows further. The tools have not changed: OpenRouter, LiteLLM, or a fifty-line router of your own. What changed is that the budget row is now Luna instead of DeepSeek, and the ultra row costs the same whichever vendor you pick.
Scorecard: what we predicted in June
We made three calls in the June edition. Two landed, one missed.
- “Fable 5 will get 30 to 50% cheaper within three months.” Missed on list price, hit on effective price. Anthropic did not cut $10 / $50. It shipped Fable 5.1 at the same list price with a 75% cheaper cache read, which is a 25% to 45% cut for agent workloads and nothing for one-shot prompts. Half credit.
- “Google will push Flash below $5 and pressure Sonnet and GPT-5.4.” Hit. Gemini 3.8 Flash is $4.50 per 1M+1M through December, and Anthropic cancelled Sonnet 5’s price rise three weeks before Google shipped it.
- “Microsoft’s in-house coding models will disrupt coding costs.” No evidence yet. Coding-cost pressure came from OpenAI’s Luna cut, not from Microsoft. Miss.
We did not predict DeepSeek raising prices. Nobody did. It is the biggest single change in this update.
BetOnAI verdict
The AI API market in September 2026 has a hard ceiling and a moving floor. The ceiling is $60 per 1M+1M and two vendors now sit on it at identical prices, which means the ultra tier will compete on cache economics and task success rate, not on the list price. The floor rose because the vendor that set it decided to charge for demand.
Three moves for anyone spending more than $500 a month:
- Re-benchmark your budget tier this week. If it is DeepSeek, check your traffic against the UTC peak window. If more than a third of your calls land in it, Luna is cheaper and simpler.
- Move mid-tier traffic to Sonnet 5 unless you have a reason not to. It is now the cheapest permanent mid-tier price from a top-three vendor.
- Do not put GPT-6 Astra or Fable 5.1 in a router yet. Measure first. On a cached agent loop Fable 5.1 wins on cost. On one-shot prompts they are identical and Opus 5 is probably good enough at half the price.
Cost per successful task, not cost per token. That rule survived the summer intact. If you want to turn the spread into a business rather than a bill, the API arbitrage play and the OpenRouter vs direct APIs margin math are the next reads.
Prices checked against vendor pricing pages on September 4, 2026. Introductory and promotional rates are marked with their end dates. This article replaces the June 2026 edition.
Sources
- Anthropic. “Pricing.” https://platform.claude.com/docs/en/about-claude/pricing
- OpenAI. “Pricing.” https://developers.openai.com/api/docs/pricing
- Google. “Gemini Developer API Pricing.” https://ai.google.dev/gemini-api/docs/pricing
- DeepSeek. “Models and Pricing.” https://api-docs.deepseek.com/quick_start/pricing
- xAI. “Models and Pricing.” https://docs.x.ai/docs/models
- Digital Applied. “AI API Pricing, August 2026: Cuts, Promos, and Traps.” https://www.digitalapplied.com/blog/ai-api-pricing-august-2026-cuts-promos-tracker
- AI and News. “DeepSeek Raises API Costs.” https://www.aiandnews.com/blog/deepseek-api-pricing-changes/
- tbreak. “Gemini 3.8 Flash is live, $0.75/$3.75 API to 31 Dec.” https://tbreak.com/gemini-3-8-flash-launch/
- AI Business. “GPT-6 Astra: What OpenAI Actually Shipped, What It Costs.” https://aibusiness.vc/tools/gpt-6-astra-launch-computer-use-pricing-what-changes-2026
How we score: read the methodology