Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Jun 1 → Jun 8, 2026·Week 24, 2026·AI & Machine Learning

Together AI raises serverless inference prices across Llama, Qwen models

Together AIPricingPrice increaseFeature added

What changed

Price increase

Multiple serverless inference models increased: Qwen3.5 9B ($0.10→$0.17 in, $0.15→$0.25 out), Llama 3.3 70B ($0.88→$1.04), Llama 3 8B Instruct Lite ($0.10→$0.14) per 1M tokens.

Feature added

NVIDIA Nemotron 3 Ultra added to serverless inference chat models at $0.60/1M input ($0.20 cached) and $3.60/1M output tokens.

Before and after

Jun 1, 2026 vs Jun 8, 2026
Pulse.
100%
Change 1 of 2 · price increased

Multiple serverless inference models increased: Qwen3.5 9B ($0.10→$0.17 in, $0.15→$0.25 out), Llama 3.3 70B ($0.88→$1.04), Llama 3 8B Instruct Lite ($0.10→$0.14) per 1M tokens.

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.