DeepInfra pricing & review
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options.
Where DeepInfra wins
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
DeepInfra model pricing
List prices refreshed 2026-08-28 ยท cheapest 25 shown| Model | Input $/1M | Output $/1M | Market floor in |
|---|---|---|---|
| Mistral Nemo | $0.019 | $0.030 | $0.019at DeepInfra |
| Llama 3.2 3B | $0.020 | $0.020 | $0.020at DeepInfra |
| Meta-Llama-3.1-8B-Instruct-Turbo | $0.020 | $0.040 | $0.020at DeepInfra |
| Llama 3.1 8B | $0.030 | $0.050 | $0.020at Nebius |
| GPT-OSS 20B | $0.030 | $0.140 | $0.015at Darkbloom |
| Llama-3-8b | $0.030 | $0.060 | $0.030at DeepInfra |
| GPT-OSS 120B | $0.037 | $0.170 | $0.030at W&B Inference |
| Qwen2.5-7B-Instruct | $0.040 | $0.100 | $0.040at DeepInfra |
| NVIDIA-Nemotron-Nano-9B-V2 | $0.040 | $0.160 | $0.040at DeepInfra |
| Llama-3.2-11b-Vision-Instruct | $0.049 | $0.049 | $0.049at Cloudflare |
| Gemma 3 12B | $0.050 | $0.150 | $0.050at DeepInfra |
| Gemma 3 4B | $0.050 | $0.100 | $0.040at AWS Bedrock |
| Mistral Small 3 | $0.050 | $0.080 | $0.050at DeepInfra |
| Nemotron 3 Nano | $0.050 | $0.200 | $0.050at Novita AI |
| Llama-Guard-3-8B | $0.055 | $0.055 | $0.020at Nebius |
| GLM-4.7 Flash | $0.060 | $0.400 | $0.060at DeepInfra |
| Ling-3.0-Flash | $0.060 | $0.180 | $0.021at OpenRouter |
| Gemma 4 26B A4B | $0.070 | $0.340 | $0.070at DeepInfra |
| Phi-4 | $0.070 | $0.140 | $0.070at DeepInfra |
| Mistral Small 3.2 24B | $0.075 | $0.200 | $0.075at DeepInfra |
| Qwen3 32B | $0.080 | $0.280 | $0.050at Lambda |
| Gemma 3 27B | $0.080 | $0.160 | $0.060at Nebius |
| Nemotron 3.5 Lightning | $0.080 | $0.200 | $0.050at OpenRouter |
| DeepSeek V4 Flash | $0.090 | $0.180 | $0.050at OpenRouter |
| Qwen3 Next 80B A3B Instruct | $0.090 | $1.10 | $0.090at DeepInfra |