Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Back
PRICING BRIEFINGDaily · Sep 12, 2026
CCerebras
Cerebras
PACKAGING CHANGE

Cerebras ditches plan cards for per-model token pricing starting at $0.35/M

Cerebras retired its three tiered plan cards and replaced them with a two-column Developer/Enterprise comparison table.

Cerebras’s pricing page · Change Detected

WHY IT MATTERS

For an inference provider whose core pitch is raw speed, abandoning packaged tiers for transparent per-token rates is a bet that price transparency, not plan structure, is what wins developer trust in a crowded GPU-inference market. Watch whether Cerebras keeps this table stable for more than a few weeks, or whether the same 10-day cadence produces another restructuring before it settles.

BY THE NUMBERS

  • Cerebras's 3rd credits/usage-billing change since Jul 20, 2026
  • 7 real changes in 90 days, ~1 every 10 days; this one just 5 days after the last
  • GPT OSS 120B: $0.35/$0.75 per M tokens; Qwen 3.8 27B: $0.99/$1.49 per M tokens
  • 596 pricing changes across 330 companies touched usage billing in 180 days (10.7%)

THEIR RATIONALE

Cerebras frames this as making its pricing transparent and usage-based rather than bundled into flat tiers, continuing a run of pricing-page iteration that has produced 7 real changes in the last 90 days at roughly one every 10 days. This lands just 5 days after Cerebras rewrote its Free Trial positioning around speed, suggesting the company is actively tuning both its message and its pricing structure in tandem rather than treating this as a one-off update.

TRACK RECORD7 changes since Jul 13, 20267 in the last 12 monthsa change every ~10 daysprevious change 5 days ago

THE SIGNAL

This is Cerebras's third move on credits and usage billing since mid-July: renaming Free to Free Trial with $5 credit framing, then a Developer Tier per-model rate table (where GPT OSS 120B briefly sold out at the same $0.35/$0.75 rate), and now a full retirement of plan cards in favor of straight per-token pricing. Usage-based credit billing is common right now — 596 pricing changes across 330 companies touched the pattern in the last 180 days — but Cerebras is cycling through it faster than almost anyone in that group, treating its pricing page as a live draft rather than a settled structure.

CREDITS & USAGE BILLING330 companies in the last 6 months10.7% of 5,546 tracked changes3× on Cerebras's pageHubSpot most often (28)

IN CONTEXT

Cerebras has touched credits and usage billing three times in under two months — renaming Free to Free Trial with $5 credit framing in July renaming Free to Free Trial, then revealing a per-model rate table a per-model rate table, and now retiring its plan cards for a straight Developer/Enterprise comparison. That's the fastest pace in the index: seven real changes in under two months, this one landing five days after Cerebras rewrote its Free Trial pitch rewrote its Free Trial pitch. It joins roughly 330 companies — including HubSpot and Datadog — that shifted pricing toward usage-based credits in the last six months, though few touch the pattern this often. For a speed-focused inference provider, ditching tiers for raw per-token rates bets that transparency, not packaging, is now the differentiator.

WHAT PEOPLE ARE SAYING

@icesmurftwitter

@cerebras is fast, but I burned through $5 USD in about 6 minutes, adding a small feature to an already written web app. 6minutes. (no it didn't finish the task) - i'm not convinced the pricing is quite right for the output. neat toy though if you wanna burn $$

View original
@PatrickMoorheadtwitter

I have always wanted to test out Cerebras for speed. So I signed up with their cloud on a paid plan. The demo showed 2,016 tokens/sec with Qwen 3.8 27B. Hermes reported 44. That's not a controlled comparison, but I kept hitting rate limits and waiting a minute to continue. I'm disappointed with the experience so far. What am I doing wrong here?

View original
@Lonly_66twitter

Qwen3.8-27B 上了 Cerebras,速度拉满:约 1,850 tokens/s,输入 $0.99/M、输出 $1.49/M。64k 上下文免费,128k 收费,Artificial Analysis 智能指数给到 34。 1,850 tok/s 是什么概念 —— 写代码补全几乎无感延迟,跑 agent 循环能快一轮是一轮。 开源模型的竞争早就从 "谁更聪明" 卷到 "谁跑得快还便宜" 了。 #Qwen #Cerebras #AI #vibecoding

View original