AWS Storage Blog
How Snap optimizes AI storage with Amazon FSx for Lustre Intelligent-Tiering
At Snap Inc., the company behind Snapchat, cost is an important factor in every infrastructure decision. With research scientists and machine learning (ML) engineers working across a wide gamut of projects, from large language models to computer vision to video generation, the AI/ML platform team faces a constant balancing act: deliver the high-performance storage that data-hungry training jobs demand, while optimizing storage spend.
In 2023, Snap adopted Amazon FSx for Lustre within an Amazon Elastic Compute Cloud (Amazon EC2) UltraCluster environment and immediately saw positive results: GPU utilization improved, job runtimes shortened, and the team serviced more high-performance workloads without adding GPUs. When Amazon Web Services (AWS) released Amazon FSx for Lustre Intelligent-Tiering with automatic data movement across three access tiers, Snap saw an opportunity to keep hot data on the fastest storage and move older, colder data to dramatically cheaper tiers, all without manual intervention.
In this post, we explore how Snap uses FSx for Lustre Intelligent-Tiering to double their file storage capacity while keeping costs flat, free up engineering resources, improve operational posture, and continue to power some of the most demanding generative AI training workloads in production today.
The challenge: Petabit-scale datasets, tight budgets, and data sprawl
Snap’s AI/ML platform supports a diverse set of generative AI workloads, including:
- Computer vision and generative model development enabling real-time camera and AR experiences
- Large-scale multimodal model training for content understanding
- Batch data processing for millions of videos
These workloads run on EC2 UltraClusters managed by Amazon Elastic Kubernetes Service (Amazon EKS) and equipped with P5 instances and P4d/P4de instances. UltraClusters provide the petabit-scale, non-blocking Elastic Fabric Adapter (EFA) networking and co-located compute needed for distributed training at scale. FSx for Lustre is able to use the EFA networking to enable throughput of 1,200 Gbps per client.
The following diagram illustrates Snap’s AI training architecture before Intelligent-Tiering.

Some of these training jobs require massive amounts of data loaded at extremely high throughput, terabytes of training data that need to be streamed to hundreds of GPUs. For these workloads, FSx for Lustre is the clear choice, delivering the throughput needed to keep expensive GPUs fully utilized.
However, not every workload needs that level of performance. Snap also uses Amazon Simple Storage Service (Amazon S3) using the Mountpoint for Amazon S3 CSI driver and the Amazon S3 Connector for PyTorch for workloads where this stack provides adequate performance. This gives the team a flexible, tiered storage strategy that matches the right storage to each workload’s requirements.
But even with this approach, a persistent challenge remained: data sprawl. With hundreds of research scientists and engineers running experiments, datasets from paused projects and changing priorities accumulated on the high-performance FSx for Lustre file systems, adding up to hundreds of terabytes of stale and cold data consuming expensive Tier 1 storage capacity. The engineering team spent valuable time trying to identify and clean up unused data, a never-ending task.
The solution: Amazon FSx for Lustre Intelligent-Tiering
Snap adopted FSx for Lustre Intelligent-Tiering, a storage class that combines the high performance of Lustre with automatic, intelligent data tiering, delivering the lowest-cost Lustre file storage in the cloud with a starting price of less than $0.005 per GB-month.
With Intelligent-Tiering, Snap’s data automatically moves across three access tiers based on different access patterns. The following table summarizes the tiers, access patterns, and benefits.
| Tier | Description | When data moves | Cost impact |
| Frequent Access | Hot data actively being used | Data accessed within the last 30 days | Baseline pricing |
| Infrequent Access | Warm data not recently accessed | Data not accessed for 30 consecutive days | Snap observed 44% cost reduction compared to Frequent Access |
| Archive Instant Access | Cold data from old projects | Data not accessed for 90 consecutive days | Up to 96% lower storage costs for infrequently accessed data compared to other managed Lustre options |
Key capabilities that made this solution ideal for Snap include:
- Fully elastic storage: The file system grows automatically as data is added. As data is deleted, you pay only for what you store. There is a maximum logical capacity of 0.5 PB for every 4 GBps of throughput provisioned.
- Automatic cost optimization: Data moves between tiers based on access patterns with no manual intervention. When data is accessed again, it automatically moves back to the Frequent Access tier.
- Same simple file system interface: Researchers and engineers interact with the same familiar POSIX-compliant file system. No application changes, no new APIs, no workflow disruptions.
- Optional SSD read cache: For latency-sensitive workloads, an optional dynamically scalable SSD read cache delivers sub-millisecond latency for frequently accessed data at HDD-level pricing. Snap enables this by default, providing an SSD cache proportional to their selected 4,000 MBps of throughput capacity.
- High performance at scale: Intelligent-Tiering file systems scale to multiple TBps of throughput and millions of IOPS, with performance that is independent of storage capacity.
Results and benefits
Snap saw several benefits from this solution.
Doubled storage capacity while maintaining costs
After adopting FSx for Lustre Intelligent-Tiering, Snap achieved a twofold expansion in file storage capacity with a flat cost structure. Snap saw inactive data, artifacts from completed experiments, departed researchers, and abandoned projects automatically tiered to lower-cost storage. Data that hadn’t been touched in months moved to the Archive Instant Access tier at a lower storage rate, while actively used training data remained in the Frequent Access tier with full performance.
“Fire and forget” simplicity
One of the most impactful benefits was operational simplicity. Snap found the Intelligent-Tiering storage class required minimal setup and no ongoing management.
There are no policies to configure, no data lifecycle rules to write, and no scheduled cleanup jobs to maintain. The tiering happens automatically based on access patterns, and data is instantly retrievable from any tier when a researcher needs it again.
Freed up engineering resources
Before Intelligent-Tiering, the platform engineering team spent significant time monitoring storage usage, identifying stale data, coordinating with researchers to confirm what could be deleted, and manually cleaning up file systems. This was a constant operational burden.
With Intelligent-Tiering handling data lifecycle management automatically, the team was able to redirect time spent on managing data and data placement toward higher-value work.
“It was an easy solution, fire and forget it, and it just works. We don’t even need to think about FSx for Lustre storage usage now.”
-Paul Roberts at Snap
Performance at scale
Despite the cost optimizations, there was no compromise on performance for active workloads. Snap continues to get the throughput and low latency their most demanding generative AI training jobs require from FSx for Lustre Intelligent-Tiering:
- Multiple GBps of throughput to keep GPU clusters fed with training data
- Millions of IOPS for cache-friendly ML workloads
- Sub-millisecond latencies for frequently accessed data through the SSD read cache
- Up to 70 percent better price-performance compared to other cloud-based Lustre storage
The file system supports a wide range of GPU instance types, from P4d instances with 400 Gbps EFA networking to P6 instances with 6,400 Gbps EFA networking, ensuring that storage performance scales alongside compute.
Solution overview
These performance capabilities are part of a broader storage architecture that Snap designed to balance throughput, cost, and operational simplicity. The following diagram illustrates the solution architecture.

The architecture follows a workload-aware storage strategy:
- FSx for Lustre Intelligent-Tiering serves as the primary storage for workloads that require the highest throughput and lowest latency. Large-scale distributed training jobs where GPU utilization depends on storage performance run on FSx for Lustre. With the Intelligent-Tiering storage class, only actively used data incurs premium storage costs.
- Amazon S3 (accessed through the Mountpoint for Amazon S3 CSI driver or the Amazon S3 Connector for PyTorch) handles workloads where that stack’s throughput and latency are sufficient, such as data preprocessing, smaller-scale experiments, and long-term dataset archival.
This combination gives Snap the flexibility to match each workload to the most cost-effective storage option while maintaining a simple, unified experience for researchers.
Technical and cost advantages
As Snap’s platform team evaluated their storage options, several technical and cost characteristics of FSx for Lustre Intelligent-Tiering stood out:
- Lowest-cost Lustre file storage in the cloud: Storage costs start at less than $0.005 per GB-month, with up to 96 percent lower storage costs for infrequently accessed data compared to other managed Lustre options.
- Up to 70% better price-performance: Snap saw up to 70 percent better price-performance compared to other cloud-based Lustre storage.
- Fully elastic: Snap pays only for what they use.
- Scales to multiple TBps and millions of IOPS: Performance is independent of storage capacity, so the team can provision the throughput they need without over-provisioning storage.
- Simple implementation: Unlike third-party storage solutions that require complex setup, ongoing management, and specialized expertise, FSx for Lustre is a fully managed service. You create a file system, mount it, and start using it.
- Compatible with existing ML frameworks: The solution works with PyTorch, TensorFlow, and other popular frameworks without modification, so Snap’s researchers interact with a standard POSIX file system.
Conclusion
Snap’s experience demonstrates a pattern that many organizations running AI/ML workloads at scale will recognize: the need for high-performance POSIX-compliant storage is growing, but so is the need to control costs as datasets grow and teams scale. FSx for Lustre Intelligent-Tiering addresses both needs simultaneously, delivering the performance that GPU-intensive training jobs demand while automatically optimizing costs for data that sits idle at any given time.
For Snap, the results speak for themselves: twice as large file storage capacity with a flat cost structure, hundreds of terabytes of cold data automatically tiered to archive storage at up to 96 percent lower cost, and multiple TBps of throughput sustained across P5 and P4d GPU clusters. The platform team eliminated manual storage cleanup entirely, redirecting that time toward higher-value work.
As Snap continues to scale their AI infrastructure, including plans to adopt next-generation P6 instances, FSx for Lustre Intelligent-Tiering provides a storage foundation that scales with them, automatically adapting to changing data access patterns and keeping costs optimized without any manual intervention.
Get started
Ready to optimize your AI/ML training storage costs? Here’s how to get started:
- Learn more about Amazon FSx for Lustre Intelligent-Tiering and its capabilities
- Read the documentation on Intelligent-Tiering storage class performance characteristics
- See another example of how FSx for Lustre Intelligent-Tiering supports SAS Grid migration
- Explore pricing on the FSx for Lustre pricing page
- Deploy a proof of concept with the Amazon FSx for Lustre Solutions Library PoC Guide
- Contact your AWS account team to discuss how FSx for Lustre Intelligent-Tiering can optimize your AI training storage infrastructure