Fireworks Announces Major Price Decrease on H100 GPU, Now Just $4.00/hour (31% Off)
What changed
Prompt caching discounts now available for enterprise deployments
FLUX.1 Kontext Max image model added at $0.08 per image
OpenAI gpt OSS 20b model added at $0.07 input, $0.30 output per 1M tokens
B200 180 GB GPU price decreased from $11.99/hour to $9.00/hour (25% decrease)
OpenAI gpt OSS 120b model added at $0.15 input, $0.60 output per 1M tokens
FLUX.1 Kontext Pro image model added at $0.04 per image
Batch inference pricing added at 50% discount of serverless pricing for input and output tokens
VLM supervised fine-tuning (fine-tuning with images) added, billed per token
Streaming ASR v2 model added at $0.0035 per audio minute
Kimi K2 Instruct model added at $0.60 input, $2.50 output per 1M tokens
Qwen3 Coder 480B model added at $0.45 input, $1.80 output per 1M tokens
Fine tuning pricing expanded from 3 tiers to 4 tiers: split DeepSeek R1/V3 ($10.00) into Models 80B-300B ($6.00) and Models >300B ($10.00)
H200 141 GB GPU price decreased from $6.99/hour to $6.00/hour (14% decrease)
H100 80 GB GPU price decreased from $5.80/hour to $4.00/hour (31% decrease)
Before and after
Jul 6, 2025 vs Oct 5, 2025Prompt caching discounts now available for enterprise deployments
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Fireworks
Full history →- Week 40, 2026·packagingFireworks adds GLM 5.3 Flash with 200K context, per-token training pricing
- Week 37, 2026·improvedFireworks rebuilds Serverless Training API catalog: drops Qwen 3.5 9B, adds DeepSeek V4 Flash and Muse Glimmer 30B
- Week 34, 2026·increaseFireworks: On-Demand GPU pricing up 11-30% from Sep 1; region surcharge to 1.5x
- Week 33, 2026·improvedFireworks: New GB300 GPU tier ($18/hr) + region premium pricing
- Week 32, 2026·packagingFireworks: New Serverless Training API added (Qwen 3.5 9B, 3.6 27B, Kimi K3)
