AI Model Pricing Comparison: OpenAI vs Claude vs Gemini
Built by R.K., Creator & Business Economics Analyst · Updated September 25, 2026Which AI provider costs less? There’s no single answer — it depends on the model, your input/output token mix, whether you use batch processing, and your cache-hit rate, and providers update pricing often enough that a fixed comparison goes stale within months. As a snapshot: OpenAI’s cheapest current model (GPT-5.6 Luna, $0.20/$1.20 per million tokens) and Google’s newest (Gemini 3.8 Flash, $0.375/$1.88) currently undercut Anthropic’s cheapest (Claude Haiku 4.5, $1/$5) on raw token rate. At the flagship tier, GPT-6 Astra and Claude Fable 5.1 both currently list $10/$50. Batch tiers typically cut all three providers’ rates by roughly 50%, and prompt caching can cut repeated input cost substantially — the exact discount varies by provider and model. Use the calculator below with your actual workload rather than relying on a fixed ranking.
Picking an AI model without running the actual cost math gets expensive fast, and the “right” number depends entirely on your own input/output token mix — a generic per-query estimate can easily be off by several multiples in either direction. The right choice also depends on your quality floor, your workload type, and which cost levers (batch, caching, model routing) you’re willing to use.
This AI model pricing comparison calculator runs your actual workload against OpenAI, Anthropic, and Google side by side, using each provider’s currently published rates. Pick a workload preset or set your own, choose a model for each provider, and the cost difference shows up immediately — at your current volume and at scale.
AI pricing changes often, sometimes monthly, and providers periodically retire or rename models. The rates built into this comparison were checked against each provider’s official pricing documentation as of late September 2026 — check the disclaimer below for direct links to each provider’s live pricing page before making a budget decision.
How to Use This AI Model Pricing Comparison Calculator
Pick a workload
Choose a preset close to your use case, or skip straight to entering your own numbers below.
Set your volume
Enter monthly runs and your average input/output token counts per run.
Choose a model per provider
Pick the specific OpenAI, Anthropic, and Google model you’d actually use for this task.
Toggle batch & caching
See how an async tier or cached prompts change the comparison before you commit to a provider.
| Volume | OpenAI | Anthropic | Lowest Estimated Cost | |
|---|---|---|---|---|
| 10,000 | $0 | $0 | $0 | — |
| 100,000 | $0 | $0 | $0 | — |
| 1,000,000 | $0 | $0 | $0 | — |
The Three Cost Levers Behind Any AI Model Pricing Comparison
The price gap between providers moves based on model tier, workload type, and which levers you’ve turned on. A team on a flagship-tier model without caching can pay several times more per run than a team running the equivalent task on a lower-cost tier from any provider. The number that matters is the specific model versus specific model comparison at your actual token counts — which is what the calculator above computes.
Model Routing, Batch Tiers, and Prompt Caching
Model routing sends classification and short responses to the cheapest current-generation tier (roughly $0.20–$0.40 input per million tokens across providers as of late September 2026) and reserves flagship models for tasks that need the depth. Batch or async tiers commonly cut token rates by roughly half for workloads that can tolerate slower turnaround, on supported models. Prompt/context caching can reduce repeated input cost substantially, but the discount and its interaction with batch pricing differ by provider — Anthropic’s cache and batch discounts stack, while Google’s batch discount does not further reduce the cached-token rate. Check your specific provider’s current documentation before budgeting around a combined discount.
How This Comparison Is Built
Model prices used in this calculator are checked against OpenAI, Anthropic, and Google’s official pricing documentation as of late September 2026 and re-checked periodically. The calculator estimates token costs for the selected models and workload, including long-context pricing tiers where a provider documents one (Gemini 3.1 Pro above 200K input tokens, GPT-6 Astra above 272K).
Cached-cost figures represent steady-state cache hits only and exclude cache-write and storage charges, which some providers charge separately. Batch pricing, long-context tiers, tool/grounding usage, taxes, credits, enterprise discounts, and other account-specific charges can all affect an actual invoice beyond what this tool models. This is a planning aid for comparing relative cost, not a guarantee of your actual invoice.
Built and verified by R.K., Creator & Business Economics Analyst
Related Tools on Ultimate Info Guide
Frequently Asked Questions
Is OpenAI cheaper than Claude in 2026?
It depends on the tier and changes as providers update pricing — there’s no fixed answer. As of late September 2026, OpenAI’s cheapest current model (GPT-5.6 Luna, $0.20/$1.20 per million tokens) undercuts Anthropic’s cheapest (Claude Haiku 4.5, $1/$5). At the flagship tier, GPT-6 Astra ($10/$50) and Claude Fable 5.1 ($10/$50) currently list the same headline rate. Use the calculator above with your actual token counts and current model selections rather than relying on a fixed comparison.
Is Gemini cheaper than OpenAI and Claude?
Often at the low end. Gemini 3.8 Flash, Google’s newest model as of September 2026, runs $0.375/$1.88 per million tokens — below GPT-5.6 Luna and Claude Haiku 4.5 on a straight per-token basis. At the high-context end, Gemini 3.1 Pro’s price roughly doubles above a 200K-token prompt ($2/$12 below the threshold, $4/$18 above), a tiering structure OpenAI and Anthropic’s current models don’t use the same way. Compare your specific token profile in the calculator rather than assuming one provider is always cheaper.
How much does batch or asynchronous pricing save?
OpenAI, Anthropic, and Google all currently document a roughly 50% token-price reduction for supported batch/asynchronous workloads, typically with a longer turnaround (Google and Anthropic both cite up to 24 hours). Content pipelines, analytics, and bulk generation are typical fits. Google’s batch discount does not additionally discount cached tokens beyond the standard cache-hit rate — Anthropic’s batch and cache discounts do stack.
How much does prompt caching save on AI costs?
It varies significantly by provider and model, and this estimate covers only steady-state cache hits, not cache-write or storage costs. Anthropic’s current cache-read discount ranges from about 90% off (Haiku 4.5, Sonnet 5) to roughly 97.5% off (Fable 5.1) versus standard input price. Google’s context-cache discount is commonly around 90% off input price for supported models. Check your specific model’s current cache-hit and cache-write rates before budgeting, since cache economics differ meaningfully from simple input/output pricing.