Together AI overhauls AI catalog with 3 new models, cuts legacy lineup
What changed
Removed legacy Llama models from Serverless Inference catalog: Llama 3 70B Instruct Reference ($0.88/$0.88 per 1M tokens) and LLaMA-2 ($0.90/$0.90 per 1M tokens).
Removed M2-BERT 80M 32K Retrieval ($0.01/1M tokens) embedding model and Salesforce LlamaRank ($0.10/1M tokens) reranking model from catalog.
Added 3 new text models: MiniMax M2.5 ($0.30/$1.20 per 1M tokens), Qwen3.5-397B-A17B ($0.60/$3.60 per 1M tokens), and GLM-5 ($1.00/$3.20 per 1M tokens) to Serverless Inference catalog.
Removed third-party model families from Serverless Inference: Mistral Instruct ($0.20/$0.20), Arcee AI Coder-Large/Maestro/Virtuoso-Large, Cogito v2 preview series (109B/405B/671B/70B), and Refuel LLM-2/LLM-2 Small.
Removed Qwen2.5 model family from Serverless Inference catalog: Qwen3 235B A22B FP8 Throughput ($0.20/$0.60), Qwen2.5 72B ($1.20/$1.20), Qwen2.5 Coder 32B Instruct ($0.80/$0.80), Qwen2.5-VL 72B Instruct ($1.95/$8), and Qwen QwQ-32B ($1.20/$1.20).
Before and after
Feb 9, 2026 vs Feb 18, 2026Removed legacy Llama models from Serverless Inference catalog: Llama 3 70B Instruct Reference ($0.88/$0.88 per 1M tokens) and LLaMA-2 ($0.90/$0.90 per 1M tokens).
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Together AI
Full history →- Week 41, 2026·improvedTogether AI adds Tev1 4B Experimental at $0.04 input with free output
- Week 39, 2026·increaseTogether AI: Qwen3.7-Max inference price increased 25% ($2.00→$2.50)
- Week 38, 2026·increaseTogether AI: Qwen3.7-Max price hike + new DeepSeek V4.1 Flash model
- Week 34, 2026·improvedTogether AI unveils Qwen3.8-2.4T-A95B, Muse Glimmer, DeepSeek V4 Pro pricing on ...
- Week 32, 2026·improvedTogether AI: HGX H100 GPU price increase + Kimi K3 & Inkling Small models added
