Cafiyn Pulse
Together · Verified 2026-09-05

Llama 3.3 70B pricing, hosted and self-hosted.

$1.04 per million tokens on Together, flat across input and output. Groq no longer publishes a public rate for it.

Llama 3.3 70B
Together · standard tier
Verify against Together's pricing page ↗
$1.04
per 1M input tokens
$1.04
per 1M output tokens

Worth knowingGroq no longer publishes a per-token price for this model (contact sales), so this is the Together rate

A typical 1,200-in / 400-out token request costs about $0.0017, or roughly $1.66 per 1,000 requests.

Llama 3.3 70B is Meta's open-weight model, which gives you a choice closed models do not: pay a hosted per-token price, or run the weights yourself on your own hardware. Together lists it at $1.04 per million tokens flat, the same for input and output, which is unusual and makes budgeting simpler than the asymmetric pricing everywhere else.

Groq, which previously offered the cheapest published rate for this model, now shows contact sales rather than a per-token price on its models page. That is worth noting as a trend: open-weight hosting rates are drifting out of public view, which makes them harder to compare rather than easier. We quote the Together figure here because it is the one you can actually read without a sales conversation.

Best for
  • Teams that want the option to move between hosts, or in-house, without a rewrite
  • Workloads with data residency or self-hosting requirements
  • Sustained high volume, where self-hosting economics eventually beat per-token pricing
Head to head

Llama 3.3 70B vs the alternatives.

ModelInput / 1MOutput / 1MSample request
Llama 3.3 70B$1.04$1.04$0.0017
DeepSeek V4 Pro$0.66$1.98$0.0016
Gemini 3.8 Flash$0.75$3.75$0.0024
Sample request assumes 1,200 input and 400 output tokens, a typical chat exchange. Your actual ratio will differ, run your own numbers in the calculator below.
FAQ

Common questions.

How much does Llama 3.3 70B cost?

$1.04 per million tokens on Together, the same rate for input and output. Groq no longer publishes a per-token price for this model, listing contact sales instead.

Is it cheaper to self-host Llama 3.3 70B or use a hosted API?

It depends on sustained volume. Hosted pricing wins until your utilisation is high enough to keep dedicated GPUs busy most of the time, at which point fixed hardware cost beats per-token billing. The crossover is a function of your duty cycle, not your total volume, so a bursty workload favours hosted for far longer than a steady one.

More models

Browse other pricing pages.

Open the tool.

Live math against Together's current pricing plus every other provider in the roster.

Compare Llama 3.3 70B on my usage