DeepInfra vs Novita AI
Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 20 of 30 shared models (input or output price), Novita AI on 4.
DeepInfra
Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..
- Among the cheapest open-model tokens anywhere
- OpenAI-compatible API
- Dedicated GPU deployments
Novita AI
Combines a serverless LLM API (DeepSeek, Llama, Qwen at aggressive per-token prices) with GPU instances and template deployments, one of the few providers covering both sides of the hosting equation..
- Both LLM API and GPU rental under one account
- Aggressive open-weight model pricing
- Template marketplace for common stacks
Shared models, priced by both
List prices refreshed 2026-08-28 ยท input $/1M tokens| Model | DeepInfra in/out | Novita AI in/out | Cheaper input |
|---|---|---|---|
| Mistral Nemo | $0.019 / $0.030 | $0.040 / $0.170 | DeepInfra |
| Llama 3.2 3B | $0.020 / $0.020 | $0.030 / $0.050 | DeepInfra |
| Llama 3.1 8B | $0.030 / $0.050 | $0.020 / $0.050 | Novita AI |
| GPT-OSS 20B | $0.030 / $0.140 | $0.040 / $0.150 | DeepInfra |
| Llama-3-8b | $0.030 / $0.060 | $0.040 / $0.040 | DeepInfra |
| GPT-OSS 120B | $0.037 / $0.170 | $0.050 / $0.250 | DeepInfra |
| Qwen2.5-7B-Instruct | $0.040 / $0.100 | $0.070 / $0.070 | DeepInfra |
| Gemma 3 12B | $0.050 / $0.150 | $0.050 / $0.100 | tie |
| Nemotron 3 Nano | $0.050 / $0.200 | $0.050 / $0.200 | tie |
| GLM-4.7 Flash | $0.060 / $0.400 | $0.070 / $0.400 | DeepInfra |
| Ling-3.0-Flash | $0.060 / $0.180 | $0.060 / $0.180 | tie |
| Gemma 4 26B A4B | $0.070 / $0.340 | $0.130 / $0.400 | DeepInfra |
| Qwen3 32B | $0.080 / $0.280 | $0.100 / $0.450 | DeepInfra |
| Gemma 3 27B | $0.080 / $0.160 | $0.119 / $0.200 | DeepInfra |
| DeepSeek V4 Flash | $0.090 / $0.180 | $0.140 / $0.280 | DeepInfra |
| Qwen3 Next 80B A3B Instruct | $0.090 / $1.10 | $0.150 / $1.50 | DeepInfra |
| Llama 4 Scout | $0.100 / $0.300 | $0.180 / $0.590 | DeepInfra |
| Qwen3.6 35B A3B | $0.100 / $0.950 | $0.248 / $1.49 | DeepInfra |
| Qwen3 30B A3B | $0.120 / $0.500 | $0.090 / $0.450 | Novita AI |
| Gemma 4 31B | $0.130 / $0.380 | $0.140 / $0.400 | DeepInfra |
| Qwen3 Next 80B A3B Thinking | $0.140 / $1.40 | $0.150 / $1.50 | DeepInfra |
| Qwen3.5 35B A3B | $0.140 / $1.00 | $0.250 / $2.00 | DeepInfra |
| Hy3 | $0.140 / $0.580 | $0.140 / $0.580 | tie |
| Qwen3 VL 30B A3B Instruct | $0.150 / $0.600 | $0.200 / $0.700 | DeepInfra |
| Qwen3 235B A22B | $0.180 / $0.540 | $0.200 / $0.800 | DeepInfra |
| DeepSeek R1 Distill Llama 70B | $0.200 / $0.600 | $0.800 / $0.800 | DeepInfra |
| Llama 4 Maverick | $0.200 / $0.800 | $0.270 / $0.850 | DeepInfra |
| Qwen3 VL 235B A22B Instruct | $0.200 / $0.880 | $0.300 / $1.50 | DeepInfra |
| Step 3.7 Flash | $0.200 / $1.15 | $0.200 / $1.15 | tie |
| Llama 3.3 70B | $0.230 / $0.400 | $0.135 / $0.400 | Novita AI |
Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.