Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Jan 3 → Apr 1, 2025·Q1 2025·AI & Machine Learning

Groq Introduces New High-Performance Models and 50% Discount on Batch API for Paid GroqCloud Customers

GroqPricingPlan removedNew planDiscount addedFeature added

What changed

Plan removed

Llama 3 Groq 8B Tool Use Preview 8k model removed (was $0.19/M tokens for both input and output at 1250 tokens/second)

New plan

Qwen QwQ 32B (Preview) 128k model added with input price $0.29/M tokens and output price $0.39/M tokens at 400 tokens/second

Plan removed

Llama 3 Groq 70B Tool Use Preview 8k model removed (was $0.89/M tokens for both input and output at 335 tokens/second)

New plan

Qwen 2.5 Coder 32B Instruct 128k model added with input price $0.79/M tokens and output price $0.79/M tokens at 390 tokens/second

New plan

DeepSeek R1 Distill Llama 70B model added with input price $0.75/M tokens and output price $0.99/M tokens at 275 tokens/second

New plan

Qwen 2.5 32B Instruct 128k model added with input price $0.79/M tokens and output price $0.79/M tokens at 200 tokens/second

New plan

Mistral Saba 24B model added with input price $0.79/M tokens and output price $0.79/M tokens at 330 tokens/second

Discount added

Batch API introduced with 25% discount rate, doubled to 50% off through end of April 2025 for paid GroqCloud customers

Feature added

Text-to-Speech (TTS) Models section introduced with PlayAI Dialog v1.0 model priced at $50.00 per million characters at 140 characters/second

Plan removed

Mixtral 8x7B Instruct 32k model removed (was $0.24/M tokens for both input and output at 575 tokens/second)

New plan

DeepSeek R1 Distill Qwen 32B 128k model added with input price $0.69/M tokens and output price $0.69/M tokens at 140 tokens/second

Plan removed

Gemma 7B 8k Instruct model removed (was $0.07/M tokens for both input and output at 950 tokens/second)

Before and after

Jan 3, 2025 vs Apr 1, 2025
Pulse.
100%
Change 1 of 12 · plan removed

Llama 3 Groq 8B Tool Use Preview 8k model removed (was $0.19/M tokens for both input and output at 1250 tokens/second)

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.