TL;DR: OpenRouter lists 427 models behind one API key and, as of September 4, 2026, adds no per-token markup: Claude Fable 5.1 is $10 in / $50 out, Claude Opus 5 is $5 / $25, Claude Sonnet 5 is $2 / $10, GPT-5.6 Terra is $2 / $12 and Gemini 3.8 Flash is $0.75 / $3.75, identical to the vendors’ pages. The cost is a 5.5% fee on card credit purchases ($0.80 minimum), 5% on crypto, and 5% on bring-your-own-key traffic only after $25,000 of list-price usage a month. Below that, BYOK is free. Eighteen models are free at 20 requests per minute. GPT-6 Astra is not listed yet.
OpenRouter’s public listing returned 427 models on September 4, 2026, including 18 free ones and 67 half-price batch variants. Every price below was checked that day against OpenRouter’s FAQ, its model listing and the vendor pages linked in the tables; the direct-vendor side is covered in our AI API pricing September 2026 roundup. The only question: what does OpenRouter cost versus direct, and where does the fee bite?
How OpenRouter Pricing Works (The Basics)
OpenRouter is pay-per-token, funded by prepaid credits, with no per-token markup: “We pass through the pricing of the underlying providers without any markup.” The money is made when you buy credits, and on BYOK volume above a generous allowance.
| Fee | Rate | Notes |
|---|---|---|
| Card (Stripe) credit purchase | 5.5%, $0.80 minimum | Charged at checkout on every top-up (OpenRouter FAQ) |
| Crypto (Coinbase) credit purchase | 5% | Never refundable |
| Per-token markup on inference | 0% | Vendor list price passed through |
| BYOK (your own vendor key) | 5% of the OpenRouter list price for that model | Only after $25,000 per month of list-price usage on pay-as-you-go; $200,000 on enterprise (BYOK docs) |
| Auto Router | 0% | “There is no additional fee for using the Auto Router” (model routing docs) |
| Opt-in prompt logging | Minus 1% | Discount on usage if you let OpenRouter log prompts |
Three things follow from that. At $100 of credits the fee is $5.50, so every “with fee” figure here is the vendor price times 1.055 (counted as a deduction from what you load, the overhead is 5.8%). The $0.80 minimum makes a $5 top-up a 16% fee, so buy in $20 to $50 blocks. And the BYOK allowance is the biggest change since June: a team spending $20,000 a month on Anthropic or OpenAI keys can route all of it through OpenRouter and pay OpenRouter nothing.
Complete Model Pricing Table (September 2026)
Prices per 1M tokens, input / output. “Direct” is the vendor’s own pricing page. The last column is 1M input plus 1M output through OpenRouter after the 5.5% card fee.
Anthropic (Claude)
| Model | OpenRouter | Direct | Match? | 1M in + 1M out, with fee |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 / $50 | $10 / $50 (Anthropic) | Yes | $63.30 |
| Claude Opus 5 | $5 / $25 | $5 / $25 | Yes | $31.65 |
| Claude Sonnet 5 | $2 / $10 | $2 / $10 | Yes | $12.66 |
| Claude Haiku 4.5 | $1 / $5 | $1 / $5 | Yes | $6.33 |
Claude Fable 5.1 went live on OpenRouter on September 1, 2026, launch day, with 1M context and a $0.25 cache-read price, per its model page, which routes to Anthropic, Azure and Google Vertex at the same $10 / $50. Fable 5 and Opus 4.8 remain listed at the same prices as their successors, and every Claude model has a :batch variant at half price. For the Claude tiers in depth, see our Anthropic API pricing guide.
OpenAI
| Model | OpenRouter | Direct | Match? | 1M in + 1M out, with fee |
|---|---|---|---|---|
| GPT-6 Astra | Not listed | $10 / $50 (OpenAI) | n/a | n/a |
| GPT-5.6 Sol | $2 / $10 default endpoint; $4 / $20 on “OpenAI Fast” | $4 / $20 promo through Nov 21, 2026 | OpenRouter lower at check time; verify | $12.66 |
| GPT-5.6 Terra | $2 / $12 | $2 / $12 | Yes | $14.77 |
| GPT-5.6 Luna | $0.20 / $1.20 | $0.20 / $1.20 | Yes | $1.48 |
| GPT-5.5 | $5 / $30 | $5 / $30 | Yes | $36.93 |
| GPT-5.4 | $2.50 / $15 | $2.50 / $15 | Yes | $18.46 |
| GPT-5.4 mini | $0.75 / $4.50 | $0.75 / $4.50 | Yes | $5.54 |
GPT-6 Astra, announced September 3, was absent from OpenRouter’s listing on September 4; OpenAI says API access is still rolling out. The GPT-5.6 Sol page carried a “50% off” tag with the default OpenAI endpoint at $2 / $10 and Flex at $1 / $5, half OpenAI’s own promo price, while Azure routes at $5 / $30 and Bedrock at $4.40 / $22. Pin a cloud provider and you pay more than direct.
| Model | OpenRouter | Direct | Match? | 1M in + 1M out, with fee |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 | $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50 (Google) | Yes | $4.75 |
| Gemini 3.5 Flash | $1.50 / $9 | $1.50 / $9 | Yes | $11.08 |
| Gemini 3.1 Pro (Preview) | $2 / $12 | $2 / $12 under 200K | Yes | $14.77 |
Gemini 3.8 Flash launched September 2 and was on OpenRouter at the promotional price within a day; expect it to double on January 1, 2027 when Google’s does.
DeepSeek, xAI, and Open-Weight Models
| Model | OpenRouter | Direct | Match? | 1M in + 1M out, with fee |
|---|---|---|---|---|
| DeepSeek V4 Flash (0731) | $0.05 / $0.16 headline (third-party host); DeepSeek’s own endpoint $0.44 / $1.32 | $0.22 / $0.66 off-peak, $0.44 / $1.32 peak (DeepSeek) | OpenRouter cheaper on default routing | $0.22 |
| DeepSeek V4 Pro (0813) | $1.12 / $3.35 | $0.66 / $1.98 off-peak, $1.32 / $3.96 peak | Between DeepSeek’s two tiers | $4.72 |
| Grok 4.6 | $2 / $6 | $2 / $6 under 200K (xAI) | Yes | $8.44 |
| Llama 4 Maverick | $0.20 / $0.70 | No single vendor | n/a | $0.95 |
| Qwen3.7 Plus | $0.32 / $1.28 | n/a | n/a | $1.69 |
| Gemma 4 31B | $0.09 / $0.34 | Also free tier | n/a | $0.45 |
| NVIDIA Nemotron 3 Super | $0.085 / $0.40 | Also free tier | n/a | $0.51 |
DeepSeek is the interesting row. Since DeepSeek introduced peak pricing on August 16, 2026, its own endpoint on OpenRouter shows the peak rate, but the V4 Flash model page lists more than 30 hosts, with OpenInference at $0.05 / $0.16 and Relace at $0.065 / $0.18. Default routing picks the cheap ones, so OpenRouter is roughly four times cheaper than DeepSeek direct off-peak and nine times at peak. The page does not disclose quantization per host; test output quality before moving production traffic.
The Hidden Costs Nobody Talks About
1. The 5.5% fee compounds with the $0.80 minimum. A $10 top-up costs $10.80 (8%). At $10,000 a month you hand OpenRouter $550, easy to forget when per-token prices look identical to direct.
2. Credits can expire. OpenRouter “reserve[s] the right to expire unused credits after one year of purchase.” Refunds only “within twenty-four (24) hours,” never for fees, never for crypto. Do not bulk-buy a year of credits.
3. Provider pinning can cost you. The model-page price is the cheapest option under default load balancing. GPT-5.6 Sol is $2 / $10 on the default route and $5.50 / $33 pinned to Azure EU.
4. Rate limits are inherited, and free models are capped hard. Paid models have no platform-level request cap but inherit upstream limits. Free models: 20 requests per minute, 50 per day until you have bought $10 of credits, then 1,000 per day (rate-limit docs). Extra keys do not help; capacity is governed globally, and a negative balance returns 402 errors “including for free models.”
5. Cache writes are real money. Caching works through OpenRouter for Anthropic, OpenAI, Google, DeepSeek, xAI and Qwen, with reads at 0.1x input price for Anthropic and DeepSeek and 0.25x for Google and xAI (prompt caching docs). But Anthropic writes cost 1.25x (5-minute) or 2x (1-hour) and OpenAI charges 1.25x on GPT-5.6 and newer: a 50K-token system prompt on Fable 5.1 is $0.625 per write.
Rate Limits, Free Models, and the Credits System
OpenRouter listed 18 free models on September 4, 2026, no card required. The notable ones:
| Free model | Context | Paid price on OpenRouter |
|---|---|---|
| Gemma 4 31B | 262K | $0.09 / $0.34 |
| NVIDIA Nemotron 3 Super 120B | 262K | $0.085 / $0.40 |
| NVIDIA Nemotron 3 Ultra 550B | 1M | $0.625 / $3.125 |
| GLM 5.2 | 256K | $0.97 / $3.04 |
| MiniMax M3 | 1M | $0.30 / $1.20 |
| Inkling (Thinking Machines) | 1M | $1 / $4.05 |
The other 11 are smaller or niche: Gemma 4 26B, three more Nemotron variants, Inkling Small, MiniMax M2.7, two Poolside Laguna models, Cohere North Mini Code, Ling 3.0 Flash Fin, Dots 3 Note Preview and Liquid LFM 2.5. Nemotron 3 Ultra 550B free at 1M context is the standout; paid, it is $3.75 per 1M in plus 1M out.
The rules: 20 requests per minute; 50 per day on a fresh account; 1,000 per day once you have ever bought $10 of credits, the best-value unlock on the platform. Free hosts may train on inputs unless you turn off “allow training” for free models in settings, and free models can be pulled without notice. Do not ship a product on a :free endpoint.
Provider Routing, Data Policies, and Fallbacks
Per the provider routing docs, default routing sorts stable providers by price and load-balances with inverse-square weighting, so a host at one third the price is nine times more likely to be picked. Overrides:
:nitrosorts by throughput and enables premium tiers;:floorsorts by price and enables budget tiers.order,onlyandignorepin, allow-list or block hosts;allow_fallbacks: falsestops OpenRouter moving a request when the pinned host fails.max_pricesets a hard ceiling per 1M prompt and completion tokens; exceed it and the request is refused, not silently upgraded.data_collection: "deny"excludes hosts that may train on prompts;zdr: truerestricts to zero-data-retention endpoints.
On data: prompts and completions are “not logged by default,” with a 1% discount for opting in. Account settings separately permit or block routing to providers that may train on paid and free traffic. One caveat from the Fable 5.1 page: Anthropic does not allow zero data retention, so zdr: true fails for Claude. The privacy docs say OpenRouter surfaces provider policies but does not route on retention rules unless asked.
Fallback chains and the Auto Router bill at “the standard rate for whichever model is selected”: drop from Fable 5.1 to Opus 5 and you pay Opus 5 prices.
OpenRouter vs Direct: A Month’s Spend, Worked
Two illustrative months. Token volumes are assumptions; prices are from the tables above.
Single-vendor product, 40M input and 8M output on Claude Sonnet 5.
- Direct: 40 x $2 + 8 x $10 = $80 + $80 = $160.00
- OpenRouter: $160 x 1.055 = $168.80, an $8.80 overhead ($105.60 a year) for one dashboard and failover to Opus 5.
Mixed stack, three models. 20M in / 4M out on Sonnet 5, 100M in / 20M out on Gemini 3.8 Flash, 50M in / 10M out on DeepSeek V4 Flash.
- Sonnet 5: 20 x $2 + 4 x $10 = $80.00
- Gemini 3.8 Flash: 100 x $0.75 + 20 x $3.75 = $150.00
- DeepSeek V4 Flash via OpenRouter default routing: 50 x $0.05 + 10 x $0.16 = $4.10
- Subtotal $234.10, plus 5.5% = $246.98 through OpenRouter.
- Direct: $80 + $150 + DeepSeek off-peak (50 x $0.22 + 10 x $0.66 = $17.60) = $247.60. If the requests land in DeepSeek’s peak window, that line becomes $35.20 and the direct total $265.20.
The mixed stack lands at parity or better on OpenRouter because the third-party DeepSeek hosts erase the fee. Single-vendor Claude traffic pays the full 5.5%.
| Monthly inference spend | OpenRouter fee (5.5%) | Yearly fee | BYOK alternative |
|---|---|---|---|
| $300 (freelancer) | $16.50 | $198 | $0 fee under the $25,000 allowance |
| $1,500 (agency) | $82.50 | $990 | $0 fee |
| $5,000 (small enterprise) | $275 | $3,300 | $0 fee |
| $30,000 (large) | $1,650 | $19,800 | Estimate: 5% on the $5,000 above the allowance, about $250 |
The old rule was “go direct above $2,000 a month.” The BYOK allowance kills it. A RAG pipeline at 10,000 queries a day (5,000 in, 500 out) on Sonnet 5 costs 50 x $2 + 5 x $10 = $150 a day, $4,500 a month: $247.50 in fees on credits, $0 on BYOK, routing kept.
Comparing OpenRouter to Alternatives
| Feature | OpenRouter | Direct APIs | LiteLLM (self-hosted) | Portkey |
|---|---|---|---|---|
| Models | 427 listed | Per vendor | “100+ LLMs” (LiteLLM docs) | Not stated |
| Overhead | 5.5% card fee, or BYOK free to $25K/month | 0% | 0% plus hosting | Free developer plan, 10K logs/month; Production $49/month (Portkey pricing) |
| Auto-failover | Yes | No | Yes, you configure it | Yes |
| Unified billing | Yes (credits) | No | No | No, bring keys |
LiteLLM is the zero-fee option if you have an engineer to run it; Portkey is a BYOK gateway on a flat subscription. OpenRouter’s edge is prepaid credits: 427 models, one card, no vendor accounts, the on-ramp resellers use in the AI API arbitrage play, where 5.5% is noise against a 5x to 20x resale markup.
Tips for Minimizing OpenRouter Costs
- Buy $10 of credits on day one. It lifts the free-model daily cap from 50 to 1,000 requests and costs $10.80.
- Switch to BYOK as soon as you hold a vendor key. Below $25,000 a month it removes the fee entirely.
- Leave routing on default for open-weight models. DeepSeek V4 Flash is $0.05 / $0.16 on default routing and $0.44 / $1.32 pinned to DeepSeek.
- Use
:batchvariants for anything not latency-sensitive. Half price on every Claude, OpenAI and Gemini model. - Set
max_priceon every production request. It fails loudly instead of overspending. - Building a consumer app? Use OpenRouter’s OAuth flow. Users connect their own account and pay their own tokens; you carry no inference cost.
What Changed Since June 2026
- Claude Fable 5.1 (September 1) and Gemini 3.8 Flash (September 2) went live within a day of launch at list price. GPT-6 Astra (September 3) is not listed yet.
- Claude Sonnet 5 at $2 / $10 replaced Sonnet 4.6 at $3 / $15 as the default mid-tier, permanent since August 10.
- GPT-5.6 Luna fell 80% on July 30 to $0.20 / $1.20, the cheapest frontier-vendor model with 1M context.
- DeepSeek’s August 16 peak pricing flipped it from “same as direct” to “cheaper on OpenRouter” via third-party hosts.
- The BYOK allowance is $25,000 a month before the 5% fee, the most important number for anyone over $2,000 a month.
- Free models: 18, down from 26 in June, but now including 1M-context options.
- Prompt caching works across nine provider families, so “OpenRouter does not cache for you” no longer holds.
BetOnAI Verdict
OpenRouter in September 2026 is a 5.5% convenience fee that disappears if you know where to look.
- Under $500 a month or multi-model: buy credits, ignore the fee, use default routing. At $300 a month the fee is $16.50, less than an hour of managing four vendor dashboards.
- $500 to $25,000 a month, mostly one vendor: bring your own key. You keep routing, fallbacks and analytics and pay OpenRouter nothing. This replaces the old “go direct above $2,000” rule.
- Above $25,000 a month: negotiate enterprise BYOK (allowance rises to $200,000) or run LiteLLM. At $30,000 a month the 5% on the excess is an estimated $250, against $1,650 on credits.
The plan, in three steps: put $10 in to unlock 1,000 free requests a day and test Nemotron 3 Ultra free first; route DeepSeek through OpenRouter, where default routing is a quarter of the off-peak direct price; and set max_price plus a spending cap on every key before you ship. Track vendor moves on our AI Pricing Watch hub.
Frequently Asked Questions
Does OpenRouter add a markup on top of model prices?
No. Pricing is passed through “without any markup.” The fee is 5.5% on card credit purchases ($0.80 minimum), 5% on crypto, and 5% on BYOK traffic only above $25,000 of list-price usage per month.
Does OpenRouter store my prompts or responses?
Not by default; opting in to logging earns a 1% discount. Prompts still go to the upstream provider, whose policy applies. Use data_collection: "deny" or zdr: true per request; Anthropic does not offer zero data retention.
Are the free models really free?
Yes, at 20 requests per minute and 50 per day, rising to 1,000 per day once you have bought $10 of credits. Free hosts may train on inputs unless disabled in settings.
Do credits expire?
OpenRouter reserves the right to expire unused credits one year after purchase. Refunds only within 24 hours, never for fees or crypto.
What happens if a provider goes down?
Default routing skips providers with recent outages and moves to the next stable, lowest-cost host. Fallbacks bill at the rate of whichever model answers, with no extra fee.
Is there a self-hosted version of OpenRouter?
No. LiteLLM is the open-source proxy that covers the same ground at zero fee on your own infrastructure.
Can I use OpenRouter for commercial applications?
Yes. Commercial use is governed by the upstream model’s terms.
Sources
- OpenRouter FAQ (fees, credits, refunds, logging): https://openrouter.ai/docs/faq
- OpenRouter rate limits: https://openrouter.ai/docs/api-reference/limits
- OpenRouter provider routing: https://openrouter.ai/docs/features/provider-routing
- OpenRouter model routing and Auto Router: https://openrouter.ai/docs/features/model-routing
- OpenRouter prompt caching: https://openrouter.ai/docs/features/prompt-caching
- OpenRouter privacy and logging: https://openrouter.ai/docs/features/privacy-and-logging
- OpenRouter BYOK: https://openrouter.ai/docs/use-cases/byok
- OpenRouter model listing (427 models, JSON): https://openrouter.ai/api/v1/models
- OpenRouter Claude Fable 5.1 model page: https://openrouter.ai/anthropic/claude-fable-5.1
- OpenRouter GPT-5.6 Sol model page: https://openrouter.ai/openai/gpt-5.6-sol
- OpenRouter DeepSeek V4 Flash model page: https://openrouter.ai/deepseek/deepseek-v4-flash-0731
- Anthropic pricing: https://platform.claude.com/docs/en/about-claude/pricing
- OpenAI pricing: https://developers.openai.com/api/docs/pricing
- Google Gemini API pricing: https://ai.google.dev/gemini-api/docs/pricing
- DeepSeek pricing: https://api-docs.deepseek.com/quick_start/pricing
- xAI models and pricing: https://docs.x.ai/docs/models
- LiteLLM documentation: https://docs.litellm.ai/docs/
- Portkey pricing: https://portkey.ai/pricing
- BetOnAI, AI API pricing September 2026: https://betonai.net/ai-api-pricing-september-2026/
How we score: read the methodology