GPU job migration infrastructure that checkpoints, migrates, and resumes live GPU workloads to increase AI throughput and cut costs up to 80%.
Cedana provides GPU live migration infrastructure that automatically checkpoints, migrates, and resumes GPU workloads across cloud instances. This enables up to 80% cost savings, 2–10x faster time to first token, and stateful reliability for training jobs even through catastrophic GPU failures. The platform integrates with Kubernetes (supporting Kueue and Slurm for distributed training and Kserve for inference serving) and works across major cloud providers. Cedana is a product of Cedana.