Containers
Simplify AI infrastructure for AWS Trainium and Elastic Fabric Adapter with Kubernetes Dynamic Resource Allocation
As organizations scale AI workloads in containerized environments, they face the complexity of managing specialized hardware that creates friction between infrastructure teams focused on stability and machine learning (ML) practitioners focused on model performance. Kubernetes Dynamic Resource Allocation (DRA) provides the foundation to solve these problems. We built the Elastic Fabric Adapter (EFA) DRA driver in the upstream DRANET project and the Neuron DRA driver for AWS Trainium to extend these benefits to customers running AI workloads on AWS. Together, these drivers deliver a unified, topology-aware resource management experience for the full stack of AWS AI infrastructure from high-performance Remote Direct Memory Access (RDMA) networking with EFA to accelerator management with AWS Trainium.
How to rapidly scale your application with ALB on EKS (without losing traffic)
To meet user demand, dynamic HTTP-based applications require constant scaling of Kubernetes pods. For applications exposed through Kubernetes ingress objects, the AWS Application Load Balancer (ALB) distributes incoming traffic automatically across the newly scaled replicas. When Kubernetes applications scale down due to a decline in demand, certain situations will result in brief interruptions for end […]
Game DevOps made easy with AWS Game-Server CD Pipeline
This is a guest post by Anita Buehrle of Weaveworks. The biggest challenge faced by game publishers is the ability to deliver new features to players as quickly as possible. Not only do new features have to arrive quickly and reliably, but they also need to be delivered in a way that optimizes costs and […]


