DeepInfra vs Groq

Same workloads, both price lists, refreshed daily. On shared line items today: DeepInfra is cheaper on 6 of 8 shared models (input or output price), Groq on 1.

DeepInfra

Consistently the price floor for open-weight models: Llama, Qwen, DeepSeek at rock-bottom per-token rates with an OpenAI-compatible API and per-request GPU billing options..

  • Among the cheapest open-model tokens anywhere
  • OpenAI-compatible API
  • Dedicated GPU deployments

Visit DeepInfra

Groq

Custom LPU silicon serving open models at hundreds of tokens/second: the speed king for latency-sensitive apps, with simple per-token pricing..

  • Fastest tokens/sec in the market
  • Simple pricing
  • Generous free tier for prototyping

Visit Groq

Shared models, priced by both

List prices refreshed 2026-08-28 ยท input $/1M tokens
Model DeepInfra in/out Groq in/out Cheaper input
GPT-OSS 20B $0.030 / $0.140 $0.075 / $0.300 DeepInfra
GPT-OSS 120B $0.037 / $0.170 $0.150 / $0.600 DeepInfra
Qwen3 32B $0.080 / $0.280 $0.290 / $0.590 DeepInfra
Llama 4 Scout $0.100 / $0.300 $0.110 / $0.340 DeepInfra
Llama Guard 4 12B $0.180 / $0.180 $0.200 / $0.200 DeepInfra
Llama 4 Maverick $0.200 / $0.800 $0.200 / $0.600 tie
Qwen3.6 27B $0.320 / $3.20 $0.600 / $3.00 DeepInfra
Kimi K2 $0.500 / $2.00 $1.00 / $3.00 DeepInfra

Cheaper-on-count is a tally of listed prices, not a quality verdict: throughput, reliability and quantization differ between providers. Disclosure.