AI API Pricing September 2026: GPT-6 Astra and Fable 5.1 at $60, Gemini 3.8 Flash at $4.50, and DeepSeek Just Tripled

Every AI API price re-checked against vendor pages on Sep 4, 2026. GPT-6 Astra and Claude Fable 5.1 both at $10/$50, Gemini 3.8 Flash at $0.75/$3.75 until Dec 31, Sonnet 5 permanent at $2/$10, and DeepSeek up 3x with peak-hour billing. Full table, per-task costs and routing math.

Bar chart of AI API prices per 1M input plus 1M output tokens, September 2026: DeepSeek V4 Flash 0.88 dollars up to Fable 5.1 and GPT-6 Astra at 60 dollars

TL;DR: Five pricing events in five weeks. OpenAI shipped GPT-6 Astra on September 3 at $10 in / $50 out, exactly matching Anthropic’s Claude Fable 5.1, which launched September 1 at the same list price with cache reads cut 75%. Google launched Gemini 3.8 Flash on September 2 at $0.75 / $3.75 until December 31, then doubles. Anthropic made Sonnet 5’s $2 / $10 permanent and cancelled the September 1 increase. And DeepSeek, the budget benchmark everyone quoted, roughly tripled its V4 prices on August 16 and added peak-hour billing. The 143x gap we reported in June is now 68x, and it closed from the bottom, not the top.

What changed since June

If you set your model budget off our June table, most of it is now wrong. For the per-vendor deep dives see the AI Pricing Watch hub, the Anthropic API pricing page and the OpenAI vs Anthropic vs Google head-to-head.

Date Vendor Change Source
Jul 30 OpenAI GPT-5.6 Luna cut 80% to $0.20 / $1.20. Terra cut 20% to $2 / $12. OpenAI pricing
Aug 10 Anthropic Sonnet 5’s introductory $2 / $10 made permanent. Scheduled $3 / $15 increase on Sep 1 cancelled. Anthropic pricing
Aug 16 DeepSeek V4 Flash from $0.14 / $0.28 to $0.22 / $0.66 off-peak and $0.44 / $1.32 peak. V4 Pro from $0.44 / $0.87 to $0.66 / $1.98 off-peak and $1.32 / $3.96 peak. DeepSeek pricing
Sep 1 Anthropic Claude Fable 5.1 at $10 / $50, cache reads $0.25 (was $1 on Fable 5). Anthropic pricing
Sep 2 Google Gemini 3.8 Flash at $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50. Gemini pricing
Sep 3 OpenAI GPT-6 Astra at $10 / $50, cached input $1. API access rolling out over the following days. OpenAI pricing

Prices are per million tokens, input / output, and were read from the vendor pages on September 4, 2026.

The full table

All prices per 1M tokens. “1M+1M” is the cost of one million input plus one million output tokens, the same shorthand we used in June so you can compare directly.

Model Input Cached input Output 1M+1M Tier
GPT-5 nano $0.05 $0.005 $0.40 $0.45 Budget
DeepSeek V4 Flash (off-peak) $0.22 $0.007 $0.66 $0.88 Budget
GPT-5.6 Luna $0.20 $0.02 $1.20 $1.40 Budget
GPT-5.4 nano $0.20 $0.02 $1.25 $1.45 Budget
DeepSeek V4 Flash (peak) $0.44 $0.014 $1.32 $1.76 Budget
Gemini 3.1 Flash-Lite $0.25 n/a $1.50 $1.75 Budget
DeepSeek V4 Pro (off-peak) $0.66 $0.022 $1.98 $2.64 Budget
Gemini 3.5 Flash-Lite $0.30 n/a $2.50 $2.80 Budget
Grok 4.3 $1.25 n/a $2.50 $3.75 Budget
Gemini 3.8 Flash (to Dec 31) $0.75 n/a $3.75 $4.50 Budget
GPT-5.4 mini $0.75 $0.075 $4.50 $5.25 Budget
DeepSeek V4 Pro (peak) $1.32 $0.044 $3.96 $5.28 Budget
Claude Haiku 4.5 $1.00 $0.10 $5.00 $6.00 Mid
Grok 4.6 $2.00 n/a $6.00 $8.00 Mid
Gemini 3.8 Flash (from Jan 1) $1.50 n/a $7.50 $9.00 Mid
Gemini 3.5 Flash $1.50 n/a $9.00 $10.50 Mid
Claude Sonnet 5 $2.00 $0.20 $10.00 $12.00 Mid
GPT-5.6 Terra $2.00 $0.20 $12.00 $14.00 Mid
Gemini 3.1 Pro (under 200K) $2.00 n/a $12.00 $14.00 Mid
GPT-5.4 $2.50 $0.25 $15.00 $17.50 Premium
Claude Sonnet 4.6 $3.00 $0.30 $15.00 $18.00 Premium
GPT-5.6 Sol (promo to Nov 21) $4.00 $0.40 $20.00 $24.00 Premium
Claude Opus 5 $5.00 $0.50 $25.00 $30.00 Frontier
Claude Opus 4.8 $5.00 $0.50 $25.00 $30.00 Frontier
GPT-5.5 $5.00 $0.50 $30.00 $35.00 Frontier
Claude Fable 5 $10.00 $1.00 $50.00 $60.00 Ultra
Claude Fable 5.1 $10.00 $0.25 $50.00 $60.00 Ultra
GPT-6 Astra $10.00 $1.00 $50.00 $60.00 Ultra
GPT-5.5 Pro $30.00 n/a $180.00 $210.00 Ultra

Notes on the table:

  • Google does not publish a separate cached-input rate on the Gemini API pricing page, so that column is blank. Gemini 3.1 Pro and the Grok models charge roughly double for prompts over 200K tokens.
  • DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else is off-peak at half price. That is Asian business hours, so if you run US-hours batch jobs you mostly pay the off-peak rate.
  • Anthropic and OpenAI both offer 50% off for batch or flex processing. Anthropic charges 1.25x input for a five-minute cache write and 2x for a one-hour write. OpenAI cache writes run 1.25x to 1.5x.
  • Grok and the Gemini 2.5 series are listed but not discussed further. They are not price leaders in any tier.

The 68x gap, and why it shrank from the bottom

In June the spread between the cheapest usable model and the most expensive was 143x: DeepSeek V4 Flash at $0.42 per 1M+1M versus Claude Fable 5 at $60.

Today it is 68x. The top did not move. Fable 5.1 and GPT-6 Astra both landed at exactly $60. The bottom moved up, because DeepSeek’s off-peak Flash is now $0.88 and its peak rate is $1.76.

That is the single most important shift in this update. The “DeepSeek is 100x cheaper” line that anchored every cost comparison for a year no longer holds. At peak hours DeepSeek V4 Flash costs more than OpenAI’s GPT-5.6 Luna, and Luna is a first-party model from a US vendor with no off-peak fine print. Third-party marketplaces are reported to resell DeepSeek V4 Flash below list (Digital Applied’s August tracker cites $0.09 / $0.18), but that is a reseller quote, not a vendor rate, and we have not verified it.

If you built anything on the assumption that DeepSeek would stay at $0.28 output forever, re-run the numbers.

What you actually pay per task

Raw token prices mislead. Below is the list-price cost of four realistic workloads. Token assumptions are stated so you can swap in your own. All figures use standard (not cached, not batch) rates, DeepSeek at off-peak, and Gemini 3.8 Flash at the introductory rate.

Task (input / output tokens) DeepSeek V4 Flash GPT-5.6 Luna Gemini 3.8 Flash Claude Sonnet 5 GPT-5.6 Terra Claude Opus 5 GPT-5.5 Fable 5.1 or GPT-6 Astra
Chatbot reply (400 / 100) $0.00015 $0.0002 $0.0007 $0.0018 $0.0020 $0.0045 $0.0050 $0.0090
Document analysis (10K / 1K) $0.0029 $0.0032 $0.011 $0.030 $0.032 $0.075 $0.080 $0.15
Agentic coding session (80K / 20K) $0.031 $0.040 $0.14 $0.36 $0.40 $0.90 $1.00 $1.80
Codebase migration (800K / 200K) $0.31 $0.40 $1.35 $3.60 $4.00 $9.00 $10.00 $18.00

Two things to notice:

  1. The budget tier has collapsed to a rounding error. A chatbot reply on Luna or DeepSeek costs two hundredths of a cent. At that level the model’s price is irrelevant and the cost of your own infrastructure dominates. Stop optimising it.
  2. The frontier tier is where per-task math actually matters. A codebase migration on Fable 5.1 or Astra costs $18 versus $3.60 on Sonnet 5. If the frontier model finishes in one attempt and Sonnet takes three, Sonnet still wins on price ($10.80) but loses on wall-clock time and review burden. This is exactly the calculation the ultra tier is priced against, and it only pays off on tasks where a failed attempt costs more than $15 of your time.

The cache-read story nobody is pricing in

The headline rates for Fable 5.1 and GPT-6 Astra are identical. The cache rates are not.

Model Input Cached input Cache discount
Claude Fable 5.1 $10.00 $0.25 97.5%
GPT-6 Astra $10.00 $1.00 90%
Claude Fable 5 $10.00 $1.00 90%
Claude Opus 5 $5.00 $0.50 90%
GPT-5.5 $5.00 $0.50 90%

For a one-shot prompt this changes nothing. For an agent that re-reads a 100K-token context on every turn it changes everything. Fifty turns over the same 100K context:

  • Fable 5 or GPT-6 Astra: 5M cached tokens at $1 = $5.00 in cached input
  • Fable 5.1: 5M cached tokens at $0.25 = $1.25

Anthropic’s launch material claims typical agentic workloads run 25% to 45% cheaper on 5.1 than on 5 at the same list price, and the cache line is where that comes from. Whether OpenAI matches it is the thing to watch in Astra’s first pricing revision. Until then, for cache-heavy agent loops the two $60 models are not the same price. Fable 5.1 is materially cheaper.

The three tiers that matter now

Budget: under $6 per 1M+1M

Models: GPT-5.6 Luna, DeepSeek V4 Flash, Gemini 3.8 Flash (until December 31), GPT-5.4 mini, Claude Haiku 4.5.

Best for: classification, extraction, routing, first-line support, anything with a clear pattern and a high volume.

The play: Luna is the new default. At $0.20 / $1.20 it is cheaper than DeepSeek at peak, comes from a first-party US vendor, and has no off-peak calendar to manage. Gemini 3.8 Flash at $0.75 / $3.75 is the capability upgrade inside the tier, but it doubles on January 1, so do not build a 2027 budget on the introductory rate.

Mid: $8 to $18 per 1M+1M

Models: Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.4, Claude Sonnet 4.6.

Best for: production workloads with real reasoning: document analysis, code review, RAG, content generation.

The play: Sonnet 5 at $2 / $10 is now the permanent mid-tier anchor, and it is $2 cheaper on output than Terra for the same input price. Sonnet 4.6 at $3 / $15 is the same family at a 50% premium. Unless you have a compatibility reason, there is no case for 4.6 over 5 anymore.

Frontier and ultra: $30 to $60 per 1M+1M

Models: Claude Opus 5, Claude Opus 4.8, GPT-5.5, Claude Fable 5.1, GPT-6 Astra.

Best for: multi-hour autonomous coding, legal and financial analysis, anything where one wrong answer costs real money.

The play: Opus 5 and GPT-5.5 at $30 and $35 do 90% of what the $60 models do. Pay for Fable 5.1 or Astra only on tasks where you have measured that the cheaper model fails, and if the task is cache-heavy, Fable 5.1’s $0.25 cache read makes it the cheaper of the two ultra models by a wide margin.

The smart routing math, updated

Nobody running production AI should run one model. Here is the September version of the routing example from June, using a team pushing 10M tokens a day (8M input, 2M output) with an 80 / 15 / 5 split by difficulty.

Route Share Model Daily cost
Simple 80% GPT-5.6 Luna $3.20
Standard 15% Claude Sonnet 5 $5.40
Hard 5% Claude Opus 5 $4.50
Routed total $13.10 per day, about $393 per month
Everything on Sonnet 5 100% $36 per day, about $1,080 per month
Everything on Fable 5.1 or Astra 100% $180 per day, about $5,400 per month

Routing saves 93% against running the ultra model on everything and 64% against running Sonnet 5 on everything. Caching, batch and off-peak pricing would cut all three rows further. The tools have not changed: OpenRouter, LiteLLM, or a fifty-line router of your own. What changed is that the budget row is now Luna instead of DeepSeek, and the ultra row costs the same whichever vendor you pick.

Scorecard: what we predicted in June

We made three calls in the June edition. Two landed, one missed.

  1. “Fable 5 will get 30 to 50% cheaper within three months.” Missed on list price, hit on effective price. Anthropic did not cut $10 / $50. It shipped Fable 5.1 at the same list price with a 75% cheaper cache read, which is a 25% to 45% cut for agent workloads and nothing for one-shot prompts. Half credit.
  2. “Google will push Flash below $5 and pressure Sonnet and GPT-5.4.” Hit. Gemini 3.8 Flash is $4.50 per 1M+1M through December, and Anthropic cancelled Sonnet 5’s price rise three weeks before Google shipped it.
  3. “Microsoft’s in-house coding models will disrupt coding costs.” No evidence yet. Coding-cost pressure came from OpenAI’s Luna cut, not from Microsoft. Miss.

We did not predict DeepSeek raising prices. Nobody did. It is the biggest single change in this update.

BetOnAI verdict

The AI API market in September 2026 has a hard ceiling and a moving floor. The ceiling is $60 per 1M+1M and two vendors now sit on it at identical prices, which means the ultra tier will compete on cache economics and task success rate, not on the list price. The floor rose because the vendor that set it decided to charge for demand.

Three moves for anyone spending more than $500 a month:

  • Re-benchmark your budget tier this week. If it is DeepSeek, check your traffic against the UTC peak window. If more than a third of your calls land in it, Luna is cheaper and simpler.
  • Move mid-tier traffic to Sonnet 5 unless you have a reason not to. It is now the cheapest permanent mid-tier price from a top-three vendor.
  • Do not put GPT-6 Astra or Fable 5.1 in a router yet. Measure first. On a cached agent loop Fable 5.1 wins on cost. On one-shot prompts they are identical and Opus 5 is probably good enough at half the price.

Cost per successful task, not cost per token. That rule survived the summer intact. If you want to turn the spread into a business rather than a bill, the API arbitrage play and the OpenRouter vs direct APIs margin math are the next reads.

Prices checked against vendor pricing pages on September 4, 2026. Introductory and promotional rates are marked with their end dates. This article replaces the June 2026 edition.

Sources

Written by Nik Sai

BetOnAI Editorial covers AI tools, business strategies, and technology trends. We test and review AI products hands-on, providing real revenue data and honest assessments. Follow us on X @BetOnAI_net for daily AI insights.

Nik Sai

BetOnAI Editorial covers AI tools, business strategies, and technology trends. We test and review AI products hands-on, providing real revenue data and honest assessments. Follow us on X @BetOnAI_net for daily AI insights.

More from Nik Sai