Free AI Cost Calculator: Claude vs GPT-5 vs Gemini vs DeepSeek across 79 models
Enter your users per day, calls per user, tokens per call, plus audio minutes, images per month, and video seconds. Get a live projection of your monthly AI bill across 22 providers, a recommendation engine that picks a primary + fallback + cheap floor for your scale, and optional email alerts if projected cost crosses a threshold.
- 40+ text, voice, image, and video models, Anthropic, OpenAI, Google, DeepSeek, xAI, Mistral, ElevenLabs, DALL-E, Flux, Runway
- Live math on your own usage: change sliders, watch bills update in real time
- Recommendation engine picks a routing strategy for your scale
- Optional email alerts when projected monthly cost crosses your threshold
Cost guides by workload
Not sure where to start? Pick the shape closest to what you are building.
Per-model pricing
Input and output cost per million tokens for each flagship model, verified 2026-09-05.
How to read an AI cost projection.
The comparator takes your own usage numbers, users per day, calls per user, tokens per call, plus any audio, image and video volume, and runs them against every model in the roster at once. The output is a projected monthly bill per model, not a price list, because a price list cannot tell you which model is cheapest for your particular shape of work.
That shape matters more than the headline rate. Output tokens cost several times input tokens at almost every provider, so a summarisation workload with long inputs and short outputs and a generation workload with the reverse can rank models in opposite orders on identical volume. Reasoning models add a further wrinkle: their intermediate reasoning is billed as output but never shown to you, so estimating from visible response length understates the bill.
Prices come from each provider's own pricing page, never from an aggregator. Several carry conditions that per-token pricing hides: DeepSeek doubles during eight hours of every weekday, xAI doubles above 200k tokens, Perplexity adds a per-request search fee, and Google's current Flash rates are introductory with a published end date. Those caveats are recorded against each model and in the change log.
Common questions
Which providers are in the Cost Comparator?
79 models from 22 providers across five categories. Text: Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5, GPT-5 mini, GPT-5 nano, o3 (reasoning), o3-mini (reasoning), o4-mini (reasoning), GPT-4o, GPT-4o mini, Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro, DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4.6, Grok 4.3, Mistral Large, Llama 3.3 70B (Together), GPT OSS 120B (Groq), GPT OSS 20B (Groq), Sonar (web-aware), Sonar Pro (web-aware), Sonar Reasoning Pro. Text to speech: ElevenLabs Multilingual v2, ElevenLabs Turbo v2, OpenAI TTS, OpenAI TTS HD, Google Cloud TTS (Neural2), Google Cloud TTS (Studio), Deepgram Aura, Azure Neural TTS. Speech to text: Whisper (OpenAI), Deepgram Nova 2, Deepgram Nova 3, AssemblyAI Universal-2, Google Cloud STT, Azure Speech-to-Text. Image: DALL-E 3 (1024), DALL-E 3 HD, Stable Diffusion 3.5, SDXL, Flux Pro, Flux Schnell, Ideogram v2, Ideogram v3, Recraft v3. Video: Runway Gen-3 Alpha, Runway Gen-3 Turbo, Pika 1.5, Luma Dream Machine, Kling 1.5, Sora Turbo. Models a provider has moved to a legacy list stay in the roster so teams still running them get accurate numbers, and are labelled as such.
How current is the pricing?
Prices come from each provider's own pricing page. A GitHub Action re-checks every one of them against that page weekly and flags any that no longer match, and a person re-verifies the full roster monthly. Every change is published with its date and source at /cost-comparator/pricing/changes.
Can I add a custom model that is not in the roster?
Yes. Click "+ Add custom model" and enter the pricing yourself. Useful for your own fine-tunes or providers we have not added yet.
How does the recommendation engine work?
We read your usage sliders (users/day, calls per user, tokens per call) and route based on volume. Under 5k calls/month: use premium models, cost is a rounding error. 5k-100k: mid-tier workhorse + premium for hard cases + cheap fallback. Over 100k: cost-optimize hard, cache aggressively, use routing.
What can I do to reduce my AI bill?
Cache prompts and responses (30-50% savings in most apps), route by difficulty (cheap model triage, premium model for hard cases), set per-user daily budgets, use prompt engineering to cut token counts, and stream tokens for perceived speed.
Are cost alerts a thing?
Yes. Turn on alerts and we email you when your projected monthly cost crosses your threshold (default $500).