AWS Cloud Operations Blog
Introducing Amazon CloudWatch Omni: Observability for the AI Era
An AI-powered observability experience for AI agents and applications, built on open standards and delivered outside of the AWS Console. Organizations are handing agents the keys to their day-to-day operations, from resolving support tickets and managing infrastructure to approving expenses, shipping production code, and increasingly the long tail of workflows that keep the business running. […]
Root Cause Analysis with Amazon Managed Service for Prometheus and AWS DevOps Agent
Introduction Teams running Prometheus-instrumented workloads, whether on Kubernetes, Amazon EC2, containers, or on-premises servers, face a growing challenge: alert fatigue from static threshold monitoring and hours spent manually investigating performance degradations. These teams can significantly reduce time spent investigating false positive alerts and manually correlating metrics by implementing automated root cause analysis. When real issues […]
Investigate your AWS account activity in plain language with Amazon Q
AWS CloudTrail now integrates with Amazon Q in the AWS Management Console, letting you investigate your AWS account activity using plain language. CloudTrail records API activity across your AWS account for security auditing, compliance, and operational troubleshooting. Getting insights from this data has traditionally meant writing queries in Amazon Athena or Amazon CloudWatch Logs Insights, […]
Reduce MTTR with AI-driven RCA using AWS DevOps Agent and Splunk
Modern cloud-native applications generate rich telemetry across metrics, logs, and deployment histories. When performance degrades, operations teams have the data, but the challenge is correlating signals across multiple tools quickly enough to minimize customer impact. Root cause analysis remains a largely manual process dependent on institutional knowledge and operator experience. This post shows how AWS […]
How Moeve scales AWS governance with automated AWS Organization Service Control Policies
Moeve, formerly known as Cepsa, is a global integrated energy company with over 90 years of experience and more than 11,000 employees. Moeve is committed to driving Europe’s energy transition and accelerating decarbonization efforts. The company has embraced digital transformation to enhance energy efficiency, safety, and sustainability, focusing on investments in green hydrogen, second-generation biofuels, […]
Analyze Application Load Balancer Logs with Amazon CloudWatch Logs
Summary Amazon CloudWatch Logs now supports Application Load Balancer (ALB) logs as vended logs, giving you out-of-the-box visibility into the health and performance of your ALB. All three ALB log types (access, connection, and health check) are delivered as structured JSON with named fields. This allows teams to attribute 5xx errors to the load balancer […]
Use AWS DevOps Agent to triage and route AWS Health event impact
Triaging the impact of AWS Health events is one of the most repetitive jobs in cloud operations, and it is exactly the kind of work AWS DevOps Agent can take on. Scheduled maintenance, operational issues, and Trust & Safety notifications (alerts about resources that may violate the AWS Acceptable Use Policy) land in your inbox […]
This Month in AWS Observability: July 2026
Introduction July was a busy month for AWS Observability. We launched features to make the telemetry you already collect more actionable, and take the operational work of collecting it off your hands. Log analytics moved closer to action with alarms that run straight from a log query and enrichment that happens at ingestion. Application-level observability […]
Multi-Cloud Observability with Amazon CloudWatch Using Bearer Token Auth and OpenTelemetry
Organizations running serverless workloads across multiple cloud providers face a specific observability challenge. There is no persistent compute to host a telemetry collector, no sidecar to attach, and no daemon running between invocations. The standard OpenTelemetry deployment model (application to local collector to a telemetry backend) does not apply in this environment. Authentication presents an […]
A Practical Guide to Amazon CloudWatch Logs Cost Optimization
Introduction As organizations centralize logs from AWS services, on-premises infrastructure, and third-party sources into CloudWatch Logs, storage becomes the dominant cost driver. Ingestion is one-time, but storage charges accumulate for the lifetime of the data, often years when compliance mandates like PCI-DSS or HIPAA apply. This post shows you how to use built-in CloudWatch Logs […]








