Together AI pricing & review
High-performance open-weight inference (Llama, DeepSeek, Qwen) on a custom stack, plus fine-tuning and GPU clusters. Consistently among the fastest and cheapest for open models.
Where Together AI wins
- Top-tier open-model throughput
- Fine-tuning pipeline
- Dedicated endpoints and clusters
Together AI model pricing
List prices refreshed 2026-08-28 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| GPT-OSS 20B | $0.050 | $0.200 | $0.015at Darkbloom |
| DeepSeek V4 Flash | $0.140 | $0.280 | $0.050at OpenRouter |
| GPT-OSS 120B | $0.150 | $0.600 | $0.030at W&B Inference |
| Qwen3 Next 80B A3B Instruct | $0.150 | $1.50 | $0.090at DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.150 | $1.50 | $0.140at DeepInfra |
| Qwen3.5-9B | $0.170 | $0.250 | $0.100at DeepInfra |
| Llama 4 Scout | $0.180 | $0.590 | $0.050at Lambda |
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.180 | $0.180 | $0.020at DeepInfra |
| GLM-4.5 Air | $0.200 | $1.10 | $0.125at Pinstripes |
| Llama Guard 4 12B | $0.200 | $0.200 | $0.180at DeepInfra |
| Llama 4 Maverick | $0.270 | $0.850 | $0.050at Lambda |
| Gemma 4 31B | $0.280 | $0.860 | $0.090at OpenRouter |
| MiniMax M3 | $0.300 | $1.20 | $0.230at W&B Inference |
| Qwen3.7 Plus | $0.320 | $1.28 | $0.320at Together AI |
| Muse Glimmer 30B | $0.350 | $1.50 | $0.300at DeepInfra |
| GLM-4.7 | $0.450 | $2.00 | $0.400at GMI Cloud |
| Kimi K2.5 | $0.500 | $2.80 | $0.450at DeepInfra |
| Inkling Small | $0.500 | $1.20 | $0.450at DeepInfra |
| Qwen3.6 Plus | $0.500 | $3.00 | $0.325at OpenRouter |
| DeepSeek-V3.1 | $0.600 | $1.70 | $0.250at DeepInfra |
| GLM-4.6 | $0.600 | $2.20 | $0.400at OpenRouter |
| Qwen3.5 397B A17B | $0.600 | $3.60 | $0.450at DeepInfra |
| Mixtral-8x7B-Instruct-V0.1 | $0.600 | $0.600 | $0.150at Anyscale |
| Nemotron 3 Ultra | $0.600 | $3.60 | $0.500at DeepInfra |
| Qwen3 235B A22B Thinking 2507 | $0.650 | $3.00 | $0.110at OpenRouter |