Cafiyn Pulse
Last checked 2026-09-07Updated monthly

LLM API price changes, tracked.

Every AI model price change we verify, with the date it happened and a link to the provider page it came from. Including the increases that are already scheduled, which are the ones you can still do something about.

ByKarthik KumarCafiyn InnovationsUpdated

Already scheduled

Published end dates on current rates. Budget against the number that starts, not the one running now.

Changes 2026-10-23
o3-mini, o4-mini

Both shut down on 23 October 2026, announced 22 April 2026 alongside o1, o1-pro and the legacy GPT-4 snapshots. OpenAI points migrations at the GPT-5.6 family.

Changes 2027-01-01
Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash

All three Flash tiers are listed at an introductory $0.75 input and $3.75 output per million tokens, marked as running through 31 December 2026.

Changes 2027-02-26
Whisper and the gpt-4o transcribe family

whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize were deprecated on 26 August 2026 and shut down on 26 February 2027. Replacements are gpt-transcribe and gpt-live-transcribe.

Tracking since 2026-09-05

The log.

  1. Scheduled increaseOpenAI

    o3-mini, o4-mini

    Both shut down on 23 October 2026, announced 22 April 2026 alongside o1, o1-pro and the legacy GPT-4 snapshots. OpenAI points migrations at the GPT-5.6 family.

    input / 1M
    $1.10
    output / 1M
    $4.40

    So what
    Six weeks. These are reasoning models, so they tend to sit on the escalation path rather than the default one, which is exactly where a failure is least likely to be noticed in testing and most expensive in production. Grep for o3-mini and o4-mini today.

    OpenAI, deprecations
  2. CorrectionOpenAI

    o3-mini, o4-mini (our listing)

    We listed both as current with no caveat. They have had a published shutdown date since 22 April 2026. Our own weekly drift check flagged both on 5 September and we recorded it as a likely false positive.

    So what
    The tool was right and the human reading it was wrong. Both entries now carry the shutdown date. If you sized a budget here on o3-mini in the last month, you priced a model that stops answering on 23 October.

    OpenAI, deprecations
  3. Scheduled increaseOpenAI

    Whisper and the gpt-4o transcribe family

    whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize were deprecated on 26 August 2026 and shut down on 26 February 2027. Replacements are gpt-transcribe and gpt-live-transcribe.

    input / 1M
    $0.01

    So what
    Whisper has one of the largest install bases of any OpenAI endpoint, and a lot of it is glue code written once and never revisited. Five months is generous by OpenAI standards; the risk is that nobody looks until it is two weeks out.

    OpenAI, deprecations
  4. Moved to legacyOpenAI

    Videos API and Sora 2

    The Videos API and the sora-2 and sora-2-pro models are removed on 24 September 2026, announced 24 March. The deprecation notice names no replacement.

    So what
    Seventeen days, and unlike every other entry on this page there is nowhere to migrate to within OpenAI. If you ship video generation on Sora, this is a vendor change, not a version bump.

    OpenAI, deprecations
  5. Scheduled increaseGoogle

    Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash

    All three Flash tiers are listed at an introductory $0.75 input and $3.75 output per million tokens, marked as running through 31 December 2026.

    input / 1M
    $0.75
    output / 1M
    $3.75

    So what
    If you are sizing a 2027 budget on today's Flash rate you are sizing it on a promotional number with a published end date. Model the post-January figure now, or pin the workload to a tier whose price is not time-boxed.

    Google, Gemini API pricing
  6. Moved to legacyGroq and Together

    Llama 3.3 70B, Llama 3.1 405B

    Groq no longer publishes a per-token price for Llama 3.3 70B, the model page now says contact sales. Together still lists Llama 3.3 70B, at $1.04 input and $1.04 output, and no longer lists Llama 3.1 405B at all.

    input / 1M
    $0.59 $1.04
    output / 1M
    $0.79 $1.04

    So what
    Open-weight hosting prices are drifting out of public view, which makes them harder to budget than closed models rather than easier. If your cost model assumes a published Groq rate, it is now an assumption rather than a number. We show the Together rate because it is the one you can actually read.

    Groq, models
  7. Moved to legacyCohere

    Command R+

    Command R+ 08-2024 is still $2.50 input and $10 output, but Cohere now lists the whole Command family under legacy models available to existing customers.

    So what
    The price has not moved, the availability has. If you are not already a Cohere customer this is not a model you can newly adopt at these rates, so drop it from any build-versus-buy comparison you are running today.

    Cohere, pricing
  8. Newly listedPerplexity

    Sonar, Sonar Pro, Sonar Reasoning Pro

    Token prices are unchanged (Sonar Pro at $3 and $15), but Perplexity also charges a per-request search fee of $5 to $14 per 1,000 requests depending on search context size, and Deep Research bills citation and reasoning tokens separately.

    So what
    Token pricing alone understates Sonar badly. At high context, 1,000 requests carries $14 of search fees before a single token is counted, which on short exchanges can exceed the token cost outright. Model this one per request, not per million tokens.

    Perplexity, pricing
  9. Price cutxAI

    Grok 4.6

    xAI now lists Grok 4.6 at $2 input and $6 output per million tokens. Grok 4, which we had at $3 and $15, is no longer on the models page.

    input / 1M
    $3.00 $2.00
    output / 1M
    $15.00 $6.00

    So what
    The current top Grok is a third of the output price of the one it replaced. Rates double on prompts of 200k tokens or more, so long-context workloads should budget at the higher tier rather than the headline.

    xAI, models and pricing
  10. Price cutMistral

    Mistral Large

    Mistral lists Large at $0.50 input and $1.50 output per million tokens. We were carrying Mistral Large 2 at $2 and $6.

    input / 1M
    $2.00 $0.50
    output / 1M
    $6.00 $1.50

    So what
    Four times cheaper on input than the figure we published. Mistral is now one of the cheapest ways to run a capable open-weight-adjacent model, and worth re-testing if you priced it out earlier.

    Mistral, pricing
  11. Newly listedDeepSeek

    DeepSeek V4 Pro and V4 Flash

    DeepSeek now lists V4 Pro at $0.66 input and $1.98 output, and V4 Flash at $0.22 and $0.66, both off-peak. V3 and Reasoner are gone from the pricing page.

    input / 1M
    $0.66
    output / 1M
    $1.98

    So what
    DeepSeek bills on a peak and off-peak clock: peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, and the rate doubles. A batch job you can shift outside those windows literally costs half as much.

    DeepSeek, API pricing
  12. Moved to legacyGoogle

    Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite

    The entire Gemini 2.5 line now appears under "legacy models" on Google's pricing page. Prices are unchanged at $1.25/$10, $0.30/$2.50 and $0.10/$0.40.

    So what
    Legacy listing is the step before a retirement date. Nothing breaks today, but 2.5 is no longer where Google puts new capability, and a migration you plan now is cheaper than one you do on a deadline.

    Google, Gemini API pricing
  13. Newly listedAnthropic

    Claude Opus 5

    Opus 5 is listed as current at $5 input and $25 output per million tokens. Opus 4.8, at $15 and $75, has moved to the legacy list.

    input / 1M
    $15.00 $5.00
    output / 1M
    $75.00 $25.00

    So what
    The current Opus tier costs a third of the one it replaces. If your routing config still escalates to Opus 4.8 you are paying 3x for the older model, and that is a one-line change.

    Anthropic, Claude pricing
  14. Newly listedOpenAI

    GPT-6 Astra

    GPT-6 Astra is listed at $10 input and $50 output per million tokens, above GPT-5.6 Sol at $4 and $20.

    input / 1M
    $10.00
    output / 1M
    $50.00

    So what
    Astra is 2.5x Sol on both sides. Treat it the way you would treat any top tier: an escalation path for the calls that need it, not the default in your client config.

    OpenAI, API pricing
  15. CorrectionOpenAI

    GPT-5

    Our roster carried GPT-5 at $5 input and $15 output. OpenAI lists $1.25 and $10. Corrected.

    input / 1M
    $5.00 $1.25
    output / 1M
    $15.00 $10.00

    So what
    Our calculator overstated GPT-5 input cost by 4x for anyone who modelled a bill on it before 5 September 2026. If you sized a budget here and ruled GPT-5 out on price, run it again.

    OpenAI, API pricing
  16. CorrectionAnthropic

    Claude Sonnet 5

    Our roster carried Sonnet 5 at $3 input and $15 output. Anthropic lists $2 and $10. Corrected.

    input / 1M
    $3.00 $2.00
    output / 1M
    $15.00 $10.00

    So what
    A third cheaper than we were showing, on both sides. Same re-run advice: if Sonnet 5 lost a cost comparison here before 5 September 2026, the comparison was wrong.

    Anthropic, Claude pricing

How this is maintained

How often is this checked?

Monthly, against each provider's own pricing page. The last full reconciliation was 2026-09-07. Between checks, treat the linked provider page as authoritative rather than this one.

Where do these numbers come from?

Provider pricing pages and changelogs only: claude.com/pricing, developers.openai.com, and ai.google.dev. We deliberately do not use pricing aggregators or summary blogs, which are wrong often enough that citing them would make this record worthless.

What is a scheduled increase?

A price the provider has already published a future change for, with a date. Google's current Flash tiers, for example, are introductory rates that end on 31 December 2026. These are the most useful entries here because they are the only ones you can still plan around.

Why do you publish your own corrections?

Because a price index that quietly fixes its mistakes is not a price index. If a number here was wrong, anyone who made a budget decision on it deserves to see that it changed and by how much.

Do you track price cuts as well as increases?

Yes, both, plus models moving to a provider's legacy list. A legacy listing is usually the step before a retirement date, so it is a cost and migration signal even when the price itself has not moved.

Open the tool.

Two of the figures in this log were wrong on our own calculator until 5 September 2026. If you modelled a bill before then, it is worth thirty seconds to check it again.

Re-run your numbers on corrected prices