DeepInfra vs Together AI
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 30 of 30 shared models (input or output price), Together AI on 0.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Together AI
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Shared models, priced by both
List prices refreshed 2026-08-28 ยท input $/1M tokens| Model | DeepInfra in/out | Together AI in/out | Cheaper input |
|---|---|---|---|
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.020 / $0.040 | $0.180 / $0.180 | DeepInfra |
| GPT-OSS 20B | $0.030 / $0.140 | $0.050 / $0.200 | DeepInfra |
| GPT-OSS 120B | $0.037 / $0.170 | $0.150 / $0.600 | DeepInfra |
| DeepSeek V4 Flash | $0.090 / $0.180 | $0.140 / $0.280 | DeepInfra |
| Qwen3 Next 80B A3B Instruct | $0.090 / $1.10 | $0.150 / $1.50 | DeepInfra |
| Llama 4 Scout | $0.100 / $0.300 | $0.180 / $0.590 | DeepInfra |
| Qwen3.5-9B | $0.100 / $0.150 | $0.170 / $0.250 | DeepInfra |
| Llama-3.3-70B-Instruct-Turbo | $0.100 / $0.320 | $1.04 / $1.04 | DeepInfra |
| Gemma 4 31B | $0.130 / $0.380 | $0.280 / $0.860 | DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.140 / $1.40 | $0.150 / $1.50 | DeepInfra |
| Llama Guard 4 12B | $0.180 / $0.180 | $0.200 / $0.200 | DeepInfra |
| Llama 4 Maverick | $0.200 / $0.800 | $0.270 / $0.850 | DeepInfra |
| DeepSeek-V3.1 | $0.250 / $0.950 | $0.600 / $1.70 | DeepInfra |
| MiniMax M3 | $0.280 / $1.10 | $0.300 / $1.20 | DeepInfra |
| Qwen3 235B A22B Thinking 2507 | $0.300 / $2.90 | $0.650 / $3.00 | DeepInfra |
| Muse Glimmer 30B | $0.300 / $1.20 | $0.350 / $1.50 | DeepInfra |
| DeepSeek V3 | $0.320 / $0.890 | $1.25 / $1.25 | DeepInfra |
| GLM-4.7 | $0.400 / $1.75 | $0.450 / $2.00 | DeepInfra |
| Qwen3-Coder-480b-A35b-Instruct | $0.400 / $1.60 | $2.00 / $2.00 | DeepInfra |
| Mixtral-8x7B-Instruct-V0.1 | $0.400 / $0.400 | $0.600 / $0.600 | DeepInfra |
| Meta-Llama-3.1-70B-Instruct-Turbo | $0.400 / $0.400 | $0.880 / $0.880 | DeepInfra |
| Kimi K2.5 | $0.450 / $2.25 | $0.500 / $2.80 | DeepInfra |
| Qwen3.5 397B A17B | $0.450 / $3.00 | $0.600 / $3.60 | DeepInfra |
| Inkling Small | $0.450 / $1.20 | $0.500 / $1.20 | DeepInfra |
| Kimi K2 | $0.500 / $2.00 | $1.00 / $3.00 | DeepInfra |
| GLM-4.6 | $0.500 / $2.00 | $0.600 / $2.20 | DeepInfra |
| Nemotron 3 Ultra | $0.500 / $2.20 | $0.600 / $3.60 | DeepInfra |
| Kimi K2.7 Code | $0.680 / $3.40 | $0.950 / $4.00 | DeepInfra |
| DeepSeek R1 | $0.700 / $2.40 | $3.00 / $7.00 | DeepInfra |
| GLM-5.2 | $0.750 / $2.40 | $1.40 / $4.40 | DeepInfra |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.