What is Fireworks AI

Fireworks AI provides serverless and dedicated inference, fine-tuning and managed training for open and proprietary AI models.

Serverless inference

As of July 31, 2026, new accounts receive US$1 in free credits. Standard serverless base-model pricing per million tokens is US$0.10 for models below 4B parameters, US$0.20 for 4B to 16B, US$0.90 above 16B, US$0.50 for mixture-of-experts models up to 56B active parameters and US$1.20 for 56.1B to 176B. Batch inference is 50% below the standard rate. Priority and Fast tiers provide different performance and pricing.

Published embedding rates start at US$0.008 per million tokens for models up to 150M parameters and US$0.016 for 150M to 350M. The current GLM-5.1 model page lists US$1.40 input, US$0.26 cached input and US$4.40 output per million tokens.

Dedicated compute and training

On-demand H100 and H200 deployments cost US$7 per GPU hour, B200 costs US$10 and B300 costs US$12, billed per second. Managed training prices vary by model size and method; for models up to 16B, published rates per million training tokens are US$0.50 for LoRA SFT, US$1 for LoRA DPO or full SFT, and US$2 for full DPO. Reserved capacity and enterprise terms are quoted separately.

Official sources: Fireworks pricing and serverless pricing documentation.