Llama 3.3 70B pricing, hosted and self-hosted.
$1.04 per million tokens on Together, flat across input and output. Groq no longer publishes a public rate for it.
Worth knowingGroq no longer publishes a per-token price for this model (contact sales), so this is the Together rate
Llama 3.3 70B is Meta's open-weight model, which gives you a choice closed models do not: pay a hosted per-token price, or run the weights yourself on your own hardware. Together lists it at $1.04 per million tokens flat, the same for input and output, which is unusual and makes budgeting simpler than the asymmetric pricing everywhere else.
Groq, which previously offered the cheapest published rate for this model, now shows contact sales rather than a per-token price on its models page. That is worth noting as a trend: open-weight hosting rates are drifting out of public view, which makes them harder to compare rather than easier. We quote the Together figure here because it is the one you can actually read without a sales conversation.
- Teams that want the option to move between hosts, or in-house, without a rewrite
- Workloads with data residency or self-hosting requirements
- Sustained high volume, where self-hosting economics eventually beat per-token pricing
Llama 3.3 70B vs the alternatives.
| Model | Input / 1M | Output / 1M | Sample request |
|---|---|---|---|
| Llama 3.3 70B | $1.04 | $1.04 | $0.0017 |
| DeepSeek V4 Pro | $0.66 | $1.98 | $0.0016 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.0024 |
Common questions.
How much does Llama 3.3 70B cost?
$1.04 per million tokens on Together, the same rate for input and output. Groq no longer publishes a per-token price for this model, listing contact sales instead.
Is it cheaper to self-host Llama 3.3 70B or use a hosted API?
It depends on sustained volume. Hosted pricing wins until your utilisation is high enough to keep dedicated GPUs busy most of the time, at which point fixed hardware cost beats per-token billing. The crossover is a function of your duty cycle, not your total volume, so a bursty workload favours hosted for far longer than a steady one.
Browse other pricing pages.
Open the tool.
Live math against Together's current pricing plus every other provider in the roster.
Compare Llama 3.3 70B on my usage