AWS Cloud Operations Blog

Category: *Post Types

Best practices for writing AWS DevOps Agent Skills

When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the […]

Automate RCA across ServiceNow, Dynatrace and Slack with AWS DevOps Agent

If you manage production incidents, you know the drill. A ServiceNow ticket fires at 2 AM. The on-call engineer wakes up, logs in to Dynatrace, pulls traces and metrics across multiple dashboards, cross-references change records, forms a hypothesis, updates the ticket, and posts findings to Slack. The investigation takes one to three hours, and that’s […]

From 2 AM alarm to answer: Security Triage with AWS DevOps Agent

Security incident triage usually starts cold. An Amazon CloudWatch alarm fires at 2 a.m. because an AWS Identity and Access Management (IAM) access key just created another access key. The responder spends the first 30 minutes assembling evidence from AWS CloudTrail, Amazon CloudWatch Logs, and Amazon Virtual Private Cloud (Amazon VPC) Flow Logs before deciding […]

Title slide with words This Month in Observability: August-September 2026 for blog covering launches

This Month in AWS Observability: August – September 2026

Introduction August and September brought the general availability of Amazon CloudWatch Omni, an AI-first, app-centric observability experience that brings together telemetry across AWS accounts, Regions, and Azure workloads in a single space. Alarms gained warm-up periods and wall clock evaluation windows, cutting the noise that comes from startup gaps and rolling-window edge cases. Database observability […]

Feature Image_Investigate your AWS account activity in plain language with Amazon Q

Investigate your AWS account activity in plain language with Amazon Q

AWS CloudTrail now integrates with Amazon Q in the AWS Management Console, letting you investigate your AWS account activity using plain language. CloudTrail records API activity across your AWS account for security auditing, compliance, and operational troubleshooting. Getting insights from this data has traditionally meant writing queries in Amazon Athena or Amazon CloudWatch Logs Insights, […]

Feature_Image_Splunk_Devops

Reduce MTTR with AI-driven RCA using AWS DevOps Agent and Splunk

Modern cloud-native applications generate rich telemetry across metrics, logs, and deployment histories. When performance degrades, operations teams have the data, but the challenge is correlating signals across multiple tools quickly enough to minimize customer impact. Root cause analysis remains a largely manual process dependent on institutional knowledge and operator experience. This post shows how AWS […]

How Moeve scales AWS governance with automated AWS Organization Service Control Policies

Moeve, formerly known as Cepsa, is a global integrated energy company with over 90 years of experience and more than 11,000 employees. Moeve is committed to driving Europe’s energy transition and accelerating decarbonization efforts. The company has embraced digital transformation to enhance energy efficiency, safety, and sustainability, focusing on investments in green hydrogen, second-generation biofuels, […]

Analyze Application Load Balancer Logs with Amazon CloudWatch Logs

Summary ​Amazon CloudWatch Logs now supports Application Load Balancer (ALB) logs as vended logs, giving you out-of-the-box visibility into the health and performance of your ALB. All three ALB log types (access, connection, and health check) are delivered as structured JSON with named fields. This allows teams to attribute 5xx errors to the load balancer […]

Use AWS DevOps Agent to triage and route AWS Health event impact

Triaging the impact of AWS Health events is one of the most repetitive jobs in cloud operations, and it is exactly the kind of work AWS DevOps Agent can take on. Scheduled maintenance, operational issues, and Trust & Safety notifications (alerts about resources that may violate the AWS Acceptable Use Policy) land in your inbox […]

This Month in AWS Observability July 2026

This Month in AWS Observability: July 2026

Introduction July was a busy month for AWS Observability. We launched features to make the telemetry you already collect more actionable, and take the operational work of collecting it off your hands. Log analytics moved closer to action with alarms that run straight from a log query and enrichment that happens at ingestion. Application-level observability […]