Qwen3.5 397B A17B API pricing
6 providers serve Qwen3.5 397B A17B. Qwen3.5 397B A17B is an open-weight model (Apache-2.0) with 397B total parameters, 17B active per token (mixture-of-experts), up to 262K context. Prices below are per 1M tokens, cheapest input first (self-hosting math further down).
Refreshed 2026-08-28 · list prices from provider APIs| Provider | Input $/1M | Output $/1M | Cache read $/1M | Context | |
|---|---|---|---|---|---|
| DeepInfra | $0.450 | $3.00output floor | $0.220 | 262K | Use |
| OpenRouter | $0.600 | $3.60 | - | 262K | Use |
| Together AI | $0.600 | $3.60 | $0.350 | 262K | Use |
| Scaleway | $0.600 | $3.60 | - | 256K | Use |
| Tensormesh | $0.600 | $3.60 | - | 262K | |
| Novita AI | $0.600 | $3.60 | - | 262K | Use |
The cheapest way to run Qwen3.5 397B A17B
On an illustrative 70M-input / 30M-output monthly workload, today's lowest listed cost is roughly $122/month, at DeepInfra. Input-heavy, output-heavy and cache-heavy workloads can produce different winners. Use the live mix calculator below rather than combining floors from two different providers. On output-token cost alone, a busy self-hosted deployment can work out cheaper (~$0.350/1M output on 3× MI300X). See the math below. Run your own numbers in the breakeven calculator, or read the breakeven math.
Your actual monthly cost
Cache read uses the listed cache rate where available; otherwise it falls back to normal input price. Batch, write-cache, volume and negotiated discounts are excluded.
Or run it yourself: Qwen3.5 397B A17B self-hosting economics
Qwen3.5 397B A17B is open-weight (Apache-2.0), so the API price above competes with the GPU-hour market. 397B total parameters (MoE, ~17B active per token) needs roughly 546 GB VRAM at FP8 or 273 GB at INT4, KV-cache headroom included.
Related models
| Model | Cheapest in $/1M | Cheapest out $/1M | Providers |
|---|---|---|---|
| Qwen3 235B A22B | $0.071 | $0.100 | 10 |
| Qwen3 32B | $0.050 | $0.100 | 9 |
| Qwen3 30B A3B | $0.051 | $0.200 | 8 |
| Qwen3 Next 80B A3B Instruct | $0.090 | $0.900 | 7 |
| Qwen3 Next 80B A3B Thinking | $0.140 | $0.900 | 7 |
| Qwen3-Coder-480b-A35b-Instruct | $0.220 | $1.50 | 7 |
FAQ
What is the cheapest API for Qwen3.5 397B A17B?
As of 2026-08-28, the lowest input price for Qwen3.5 397B A17B is DeepInfra at $0.450 per 1M input tokens. The lowest output price is DeepInfra at $3.00 per 1M output tokens. The cheapest provider for a real workload depends on its input/output mix. The next-lowest input price is OpenRouter at $0.600, a 33% difference.
How much VRAM do you need to self-host Qwen3.5 397B A17B?
Qwen3.5 397B A17B has 397B parameters, so plan for roughly 546 GB of VRAM at FP8 or 273 GB at INT4/AWQ, KV-cache headroom included. That exceeds a single GPU. A typical node is 3× MI300X.
← All models · GPU rental prices · Breakeven calculator · Get this data as JSON