Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Packaging change·Jan 3 → Apr 9, 2025·Q1 2025·Marketing & Content

Fireworks Unveils Major 10x Increase in Serverless Inference RPM for Developer Plan, Boosting Capacity to 6,000

FireworksPackagingNew planFreemium changeThreshold addedLimit decreaseLimit increase

What changed

New plan

Meta Llama 4 Scout (Basic) model added with pricing $0.15 input, $0.60 output per 1M tokens

New plan

DeepSeek R1 (Fast) model added with pricing $3.00 input, $8.00 output per 1M tokens

New plan

Streaming transcription service added at $0.0032 per audio minute

Freemium change

Spending tier qualification thresholds reduced - Tier 2 from $100+ to $50+, Tier 3 from $1,000+ to $500+, Tier 4 from $10,000+ to $5,000+

New plan

Meta Llama 4 Maverick (Basic) model added with pricing $0.22 input, $0.88 output per 1M tokens

Threshold added

Developer plan serverless inference daily token limit added: 2.5 billion tokens/day

Threshold added

Developer plan monthly GPU hours limit added: 2,000 GPU hours/month

Limit decrease

Developer plan on-demand GPU deployment reduced from 16 GPUs to 8 GPUs

Limit increase

Developer plan serverless inference RPM increased from 600 to 6,000 (10x increase)

New plan

DeepSeek R1 (Basic) model added with pricing $0.55 input, $2.19 output per 1M tokens

Before and after

Jan 3, 2025 vs Apr 9, 2025
Pulse.
100%
Change 1 of 10 · plan added

Meta Llama 4 Scout (Basic) model added with pricing $0.15 input, $0.60 output per 1M tokens

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.