
Baseten cuts DeepSeek V4.1 Flash cache pricing 77%, its third rate move since July
Baseten cut the cache-input price for DeepSeek V4.1 Flash on its Model APIs price list from $0.03 to $0.007 per 1M tokens, a 77% reduction.
Baseten’s pricing page · Change Detected
WHY IT MATTERS
For teams running high-context, high-repeat workloads on Baseten, this cut lowers the effective cost of caching strategies specifically, which could tilt architecture decisions toward leaning harder on cached context rather than fresh input. It also signals Baseten is treating its Model APIs price list as a live lever, adjusted roughly twice a month, rather than a fixed rate card - competitors hosting the same open models will feel pressure to match on cache pricing in particular. Watch whether the next move is another cache-input cut on a different model, or a shift toward competing on output price instead.
BY THE NUMBERS
- Cache-input price down 77%: $0.03 to $0.007 per 1M tokens
- 3rd pricing move in ~2 months; 5 changes since tracking began Jul 20
- Previous cut: GLM 5.2 cache input $0.26 to $0.14/1M on Jul 27
- Credits/usage-billing moves = ~11% of all index changes in last 180 days
THEIR RATIONALE
This is Baseten's third pricing move since late July, and its second cache-input cut in that span: GLM 5.2's cache-input price fell from $0.26 to $0.14/1M tokens on July 27, and DeepSeek V4.1 Flash now gets a steeper cut. Cache-input tokens are typically priced well below fresh input, so trimming that line targets high-repeat-context workloads like agents, RAG, and chat history without touching the input/output rates that set the model's baseline margin. The cut follows Baseten's August 3 expansion of its Model APIs lineup with Kimi K3, Inkling-Small, and DeepSeek-V4-Flash-0731, suggesting a pattern of adding models first and sharpening price on the highest-volume line items after.
THE SIGNAL
Baseten has made 5 real changes to its pricing page in the roughly two months since tracking began on July 20 - about one every 15 days, split between three price cuts and two feature additions, all concentrated on the Model APIs surface rather than core compute or serving plans. That reads as an inference-cost fight being waged token type by token type: cache input is now the recurring target, while raw input and output prices hold steady. Zoomed out, pricing moves tied to credits and usage-metered billing account for roughly 11% of all changes tracked across the index in the last 180 days - 629 changes across 347 companies, led by usage-heavy platforms - evidence that granular, line-item cuts on metered inference are becoming a standard lever for AI infra vendors, not just a Baseten habit.
IN CONTEXT
This is Baseten's second cache-input price cut in under two months, following GLM 5.2's cut from $0.26 to $0.14 per 1M tokens in late July, and its third pricing move since it added Kimi K3, Inkling-Small, and DeepSeek-V4-Flash-0731 in early August. Each cut has hit the cache-input line specifically, leaving standard input and output rates untouched - a narrower version of the usage-based repricing playing out across roughly 11% of all pricing changes tracked over the last six months, from HubSpot's credit shifts to Notion's move to bill credits for its Workers beta starting in October. For a page tracked only since July, Baseten's cadence - a change roughly every 15 days - reads like active price management of a fast-growing model catalog.