Baseten adds GLM-5.3 Fast, retires 6 models, cuts DeepSeek V4.1 Flash cache price 77%
New model GLM-5.3 Fast added to the Model APIs price list at $2.10 input / $0.21 cache input / $6.60 output per 1M tokens.; DeepSeek V4.1 Flash cache-input price cut from $0.03 to $0.007 per 1M tokens (-77%); input $0.30 and output $1.20 unchanged.; GLM 4.7 ($0.60/$0.12/$2.20) and DeepSeek V4 Pro ($1.74/$0.145/$3.48 per 1M tokens) removed from the Model APIs price list; DeepSeek V4 Pro 0813 remain
What changed
DeepSeek V4.1 Flash cache-input price cut from $0.03 to $0.007 per 1M tokens (-77%); input $0.30 and output $1.20 unchanged.
Kimi K2.6 and Kimi K2.7 Code ($0.95/$0.16/$4.00), Inkling-Small ($0.50/$0.10/$1.20) and Inkling ($1.00/$0.17/$4.05 per 1M tokens) removed from the Model APIs price list; Kimi K3 remains.
New model GLM-5.3 Fast added to the Model APIs price list at $2.10 input / $0.21 cache input / $6.60 output per 1M tokens.
GLM 4.7 ($0.60/$0.12/$2.20) and DeepSeek V4 Pro ($1.74/$0.145/$3.48 per 1M tokens) removed from the Model APIs price list; DeepSeek V4 Pro 0813 remains.
Before and after
Sep 21, 2026 vs Sep 28, 2026DeepSeek V4.1 Flash cache-input price cut from $0.03 to $0.007 per 1M tokens (-77%); input $0.30 and output $1.20 unchanged.
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
