Cedana AI Compute Fabric automatically checkpoints and migrates live GPU workloads across Amazon EKS and Slurm without losing progress. Achieve up to 2x higher AI job throughput per GPU while increasing utilization, improving reliability, and enabling dynamic prioritization.
Cedana AI Compute Fabric provides system-level checkpointing and migration for GPU workloads running on Amazon EKS and Slurm.
Cedana makes execution state portable across nodes and instances, allowing AI training, fine-tuning, inference, and distributed workloads to pause, move, and resume without losing progress.
Unlike application-level checkpoints, Cedana operates transparently at the system layer, requiring no code changes while preserving full process state, GPU memory, and distributed context.
By decoupling AI workloads from fixed infrastructure, Cedana increases GPU utilization and delivers up to 2x higher AI job throughput per GPU. Workloads automatically recover from node failures, spot interruptions, and maintenance events without restarting from scratch.
Teams can dynamically reprioritize jobs, rebalance clusters, consolidate underutilized GPUs, and safely run long jobs on Spot instances.
The result:
Improved reliability
Reduced wasted compute
Lower cloud costs
Shorter queue times
Higher productivity per $/GPU
Cedana integrates in minutes with Amazon EKS, and Slurm environments and supports single-node and distributed multi-GPU/CPU workloads.
Ideal for AI startups, research labs, enterprises, and platform teams operating multi-tenant GPU clusters, Cedana enables infrastructure automation, spot resilience, SLA enforcement, and efficient AI factory operations across AWS.
Highlights
Cloud-Native GPU Checkpointing for Amazon EKS Automatically checkpoint and migrate AI workloads across Amazon EKS without code changes. Preserve full execution state, including GPU memory and distributed processes, enabling seamless recovery from node failures, spot interruptions, and autoscaling events.
Increase Throughput 2x and Reduce GPU Wait Times Boost AI training and inference throughput by eliminating lost work from failures and preemptions. Cedana improves GPU utilization, enables dynamic job prioritization on Amazon EKS and Slurm, and reduces queue times across multi-tenant GPU clusters.
Automate Spot Instances for Long-Running AI Jobs Run training and stateful inference workloads reliably on Amazon EC2 Spot Instances without losing progress. Cedana automatically checkpoints and resumes GPU workloads across interruptions, enabling resilient Spot usage, lower cloud costs, and significantly higher throughput per $/GPU on Amazon EKS.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay based on usage, measured per GiB-hour of instance memory under management. Your cost scales with two factors: how much memory your managed workloads consume and how long that memory stays under management. There are no fixed tiers or instance-size choices. The more memory you place under management and the longer it runs, the more you pay. This single usage metric covers the platform's save, migrate, and resume capability across CPU and GPU containers.
Top-of-mind questions for buyers
What counts as one GiB-hour of instance memory under management?
One GiB-hour equals one gibibyte of workload memory held under management for one hour. The platform tracks the memory footprint of your managed CPU and GPU containers, then multiplies that by the time it stays under management. Both the memory size and the runtime duration determine the count.
Am I charged when a managed workload is suspended or paused?
The platform can suspend and resume workloads based on demand and take continuous snapshots. Billing meters instance memory under management per hour. When memory is no longer under management, that memory stops accruing charges. Suspending workloads to eliminate idle resources can reduce the memory counted against your usage.
Does the price change based on whether I run CPU or GPU workloads?
No. The single usage metric is instance memory under management, measured in GiB-hours. It applies the same way across CPU and GPU containers. The platform supports save, migrate, and resume across both processor types, but billing does not vary by workload type — only by memory size and duration.
docs.cedana.ai+2
Helpful?
Vendor refund policy
Contact our support team for refund information.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
AWS seamlessly extends Nexus Dashboard capabilities to the cloud. It provides a centralized management console that allows network operators easily access applications and perform the lifecycle management of their fabric, from provisioning, troubleshooting, or gaining deeper visibility into their network
Stratio Generative AI Data Fabric is an end-to-end data management platform of choice for enterprises in banking, insurance, manufacturing, retail, and government sectors. Talk to your data in natural language, automate data discovery, data governance, and semantic business meaning.
This product has charges associated with the pre-built hardening to the CIS Benchmarks™ and recurring maintenance targeted specifically at GPU optimized AMIs. The CIS Hardened Images® are hardened in accordance with the associated CIS Benchmarks, an industry best practice for secure configuration. Reduce cost, time, and risk by building your AWS solution with CIS AMIs.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.