featureSep 11, 2026
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together AI has integrated new open-weight models and introduced granular training features including Expert LoRA, early stopping, and live experiment tracking. The update also implements pre-flight validation and tokenized dataset previews alongside reduced pricing for specific models.
➔ Service now includes Expert LoRA, live experiment tracking, early stopping, pre-flight validation, tokenized dataset previews, and reduced pricing on selected models.
featureSep 10, 2026
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Together AI has ported the ThunderKittens library to support the NVIDIA Vera Rubin NVL72 architecture. The implementation includes a rebuilt NVFP4 GEMM kernel that achieves over 22 PFLOPS performance.
➔ ThunderKittens supports NVIDIA Vera Rubin NVL72 with a rebuilt NVFP4 GEMM kernel delivering over 22 PFLOPS.
pricingSep 10, 2026
Introducing preemptible compute: the same compute, half the price
Together AI has introduced preemptible compute instances for its GPU clusters. This new offering provides identical GPU capacity at a 50% discount compared to on-demand rates, subject to a five-minute drain window.
➔ Together GPU Clusters now offer preemptible compute instances at 50% of the on-demand rate with a five-minute termination notice.
product launchSep 9, 2026
The Open Source AI Stack
Together AI has released a comprehensive open-source AI stack designed to streamline the deployment and fine-tuning of large language models. The stack integrates optimized inference engines, data processing pipelines, and model training frameworks to reduce latency and infrastructure overhead.
➔ Users have access to a unified, open-source stack for deploying, fine-tuning, and serving open-weights models with optimized performance.
featureSep 5, 2026
Together AI Rolls Out Next-Generation Platform Capabilities & API Architecture
Together AI released substantial architectural upgrades, introducing enhanced API integrations and updated workflow tooling for its Artificial Intelligence ecosystem.
➔ Modernized Together AI platform capabilities with enhanced integration tooling.
product launchAug 28, 2026
GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
Together AI has introduced GLM-5.3 Flash as a high-efficiency alternative to the standard GLM-5.3 model. The new model achieves a 17x reduction in cost while maintaining performance within 5.6 points of pass@1 and 2.6 points of pass@4 on the DeepSWE benchmark.
➔ Availability of GLM-5.3 Flash, offering a 17x cost reduction compared to GLM-5.3 with a 5.6 point pass@1 performance delta.