AWS Cloud Operations Blog
Category: Best Practices
Best practices: Implementing observability with AWS
As customers deploy cloud-based solutions, they need to be able to ensure that systems are running smoothly, and that they can quickly remediate issues when they arise. Deploying observability at scale can be challenging for customers, especially when it involves tens and hundreds of services across their enterprise. Customers want best practice recommendations, guidance in […]
How to perform a Well-Architected Framework Review- Part 3
In previous blog posts, we discussed the first two phases for running a Well-Architected Framework Review, or WAFR. The first phase is to Prepare and the second phase in to conduct the Review. In this blog post, we dive deep into the third phase: Improve. Figure-1 WAFR Phases What is the Improve phase? At this […]
How to perform a Well-Architected Framework Review- Part 2
There are three phases to conduct a successful Well-Architected Framework Review or WAFR: Prepare, Review and Improve. In part 1 of this blog series, we discussed the preparation phase. In this part, we will dive deep into the best practices of the second phase, the actual review. Figure-1 WAFR Phases Assuming you follow the recommendations […]
How to perform a Well-Architected Framework Review- Part 1
Is my workload well-architected? Is my team following cloud best practices? How do other customers implement solution X? What is the best way to configure service Y? These are examples of questions I usually get from my customers who want to validate if their architecture is aligned with AWS best practices. The answers to these […]
Manage continuous compliance by using AWS Config Configuration Recorder resource type
AWS Config recently added support for configuration recorder as a resource type. The AWS::Config::ConfigurationRecorder resource is a configuration item (CI) for configuration recorder that tracks changes to the state of AWS Config configuration recorder (configuration recorder). You can use this CI to check if the state of the configuration recorder has changed (drifted), from its […]
Optimizing alarm lifecycle with Amazon CloudWatch Metrics Insights alarms
Do you have entire fleets of dynamically changing resources that you are struggling to easily monitor and set alarm on? Do you have a ton of dangling alarms that you are paying for and that is cluttering your view? Are you looking for a simplified way to create alarms that automatically adjusts to resources that […]
Increase visibility and governance on cloud with AWS Cloud Operations services – Part 2
Introduction This blog post is a continuation of Part 1. To recap, as your organization adopts AWS, you will likely leverage multi-account architectures to meet your requirements. We introduced some foundational patterns to prepare the environments for centralized operations and governance using AWS Cloud Operations services. In this blog (Part 2), we will show you […]
Maximize Cloud Adoption Benefits with a Well-Architected Organizational Culture
Organizational culture, often described as the “personality” of an organization, determines how people work, interact, and respond to change and challenges. There is strong recognition, supported by evidence, that an organization’s culture is a powerful determinant of transformation success. Culture’s impact is magnified in cloud transformation, where the cloud’s extraordinary capabilities are limited only by […]
Migrating to Amazon Managed Service for Prometheus with the Prometheus Operator
The Prometheus Operator allows cluster administrators to manage Prometheus clusters running in Kubernetes. It makes it easy to deploy and manage Prometheus via native Kubernetes components. In this blog post, I will demonstrate how you can deploy Prometheus via the Prometheus Operator, and how you can easily migrate your monitoring workloads to take advantage of […]
Using the Fault Tolerance Analyser Tool to Identify Potential Issues
Introduction Ensuring resilience, the ability for a system to recover from a failure induced by load, attacks, and other issues, is a shared responsibility that underpins the reliability of your workloads. While AWS provides the resilient underlying cloud infrastructure, customers are tasked with maintaining the resilience of their applications. In this landscape of joint responsibility, […]









