Pricing built for growth

Production inference that won't break
your product or your bank.

Pay as you go

Serverless endpoints

Per-token access to 2M+ open models. Call an endpoint and go — zero setup.

  • 2M+ open models
  • No minimums
  • Community support
Get Started
Early access
Per million tokens

Elastic Endpoints

Private, optimized endpoints priced per million tokens.

  • Private optimized endpoints
  • Per-workload tuning
  • Dedicated support
Talk to sales
Reserved capacity

Dedicated deployments

Reserved GPUs sized to your roadmap, with negotiated SLAs.

  • Reserved GPUs, sized to you
  • Negotiated latency SLA
  • Dedicated support
Talk to sales
Lowest rate

Batch

High-throughput offline jobs at the lowest per-token rate, on spare fleet capacity.

  • Lowest per-token rate
  • Millions of requests per job
  • Great for evals & embeddings
Get Started
Elastic Endpoints

Per-token model pricing

Pay only for what you use. No minimums, no rate limits.

Price per 1M tokens
Model Input Output Cache read
Kimi K3 $3.00 $15.00 $0.30
Kimi K2.7 Code $0.75 $3.50 $0.16
Kimi K2.6 $0.75 $3.50 $0.16
GLM-5.3 Flash $0.15 $0.50 $0.03
GLM-5.3 $1.40 $4.40 $0.26
GLM-5.2 $1.40 $4.40 $0.26
GLM-5.1 $1.40 $4.40 $0.26
GLM-5 $1.00 $3.20 $0.20
MiniMax M3 $0.30 $1.20 $0.06
MiniMax M2.5 $0.30 $1.20 $0.03
DeepSeek V4 Pro $1.74 $3.48 $0.10
DeepSeek V4 Flash $0.14 $0.28 $0.07
Qwen3.6 35B-A3B $0.15 $1.00 $0.05
Qwen3.5 397B-A17B $0.50 $3.60 $0.30
Qwen3.5 35B-A3B $0.15 $1.00 $0.05
Qwen3-Coder-Next $0.12 $0.80 $0.07
Qwen3-VL 235B-A22B $0.21 $1.90 $0.10
Qwen3-VL 8B $0.25 $0.75 $0.12
Qwen3-Next 80B $0.10 $1.10 $0.07
Qwen3 235B-A22B (2507) $0.14 $0.80 $0.05
Qwen2.5-VL 72B $0.80 $1.00 $0.40
Mistral Small 3.2 24B $0.09 $0.30 $0.05
Llama 4 Maverick (FP8) $0.35 $1.00 $0.17
Llama 3.3 70B (FP8) $0.22 $0.50 $0.11
Nemotron 3 Ultra 550B (NVFP4) $0.50 $2.50 $0.10
MiMo v2.5 $0.14 $0.28 $0.05
Resemble TTS (English) $18.50
Trinity Large (Thinking) $0.22 $0.85 $0.06
Gemma 4 26B-A4B $0.13 $0.40 $0.05
Gemma 4 31B $0.15 $0.40 $0.06
Skyfall 31B v4.2 $0.55 $0.80 $0.25
Cydonia 24B v4.1 $0.30 $0.50 $0.15
gpt-oss-120b $0.10 $0.75 $0.055
gpt-oss-120b (Fast) $0.15 $0.60
gpt-oss-20b $0.04 $0.20 $0.02
UI-TARS 1.5 7B $0.10 $0.20 $0.10
Gemma 3 27B $0.08 $0.45 $0.04
Skyfall 36B v2 (FP8) $0.55 $0.80 $0.25
BGE-M3 $0.01

Start building today

Instantly run any open model — popular or specialized.