Introducing preemptible compute: the same compute, half the price
Together AI has introduced preemptible compute instances for its GPU clusters. This new offering provides identical GPU capacity at a 50% discount compared to on-demand rates, subject to a five-minute drain window.
Verified State Diff
Impact & Verification Analysis
AI researchers, machine learning engineers, and enterprise developers running fault-tolerant or batch-processed GPU workloads.
This significantly lowers the barrier to entry for large-scale model training and experimentation by reducing infrastructure costs by half, enabling more efficient resource allocation for non-latency-critical tasks.
Full Fact Overview
Together AI is expanding its infrastructure-as-a-service model by introducing preemptible instances, a common cloud-native pattern for optimizing GPU utilization. By allowing the platform to reclaim capacity with a five-minute notice, Together AI can offer lower-cost compute for fault-tolerant workloads such as distributed training, batch inference, or non-time-sensitive fine-tuning jobs. This move aligns Together AI's pricing strategy with major hyperscalers like AWS (Spot Instances) and GCP (Preemptible VMs), increasing their competitiveness for cost-sensitive enterprise AI development.