Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Nov 1 → Feb 1, 2025·Q4 2024·AI & Machine Learning

Huggingface Reduces GCP NVIDIA A100 Inference Endpoint Prices by 40% to $3.60/hour, Delivering Significant Cost Savings

HuggingfacePricingPrice decreaseFeature changeFeature addedThreshold addedLimit increase

What changed

Price decrease

TPU v5e 2x4 Spaces Hardware price decreased from $11.00/hour to $9.50/hour (14% decrease)

Feature change

Pro Account Inference API feature enhanced from 'higher rate limits' to 'x20 higher rate limits on Serverless API'

Price decrease

TPU v5e 1x1 Spaces Hardware price decreased from $1.38/hour to $1.20/hour (13% decrease)

Price decrease

GCP NVIDIA H100 Inference Endpoints prices decreased (1-GPU from $12.50/hour to $10.00/hour, 20% decrease)

Feature added

Enterprise Hub now includes '5x more ZeroGPU quota for all org members'

Threshold added

New AWS CPU Inference Endpoint tier added: 16 vCPUs / 32GB at $0.54/hour

Price decrease

GCP Intel Sapphire Rapids CPU Inference Endpoints price decreased (1 vCPU from $0.07/hour to $0.05/hour, 29% decrease)

Price decrease

GCP NVIDIA L4 1-GPU Inference Endpoints price decreased from $1.00/hour to $0.70/hour (30% decrease)

Price decrease

GCP NVIDIA A100 Inference Endpoints prices decreased significantly (1-GPU from $6.00/hour to $3.60/hour, 40% decrease)

Limit increase

Nvidia A100 - large Spaces Hardware VRAM increased from 40 GB to 80 GB at same $4.00/hour price

Price decrease

TPU v5e 2x2 Spaces Hardware price decreased from $5.50/hour to $4.75/hour (14% decrease)

Before and after

Nov 1, 2024 vs Feb 1, 2025
Pulse.
100%
Change 1 of 11 · price decreased

TPU v5e 2x4 Spaces Hardware price decreased from $11.00/hour to $9.50/hour (14% decrease)

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.