Skip to main content

Amazon CloudWatch Omni

Observability for the AI era

Amazon CloudWatch Omni

Amazon CloudWatch Omni is the next evolution of Amazon CloudWatch: a re-imagined observability experience that is AI-powered, built on open standards, and delivered outside the AWS Console. Purpose-built for the AI era, CloudWatch Omni unifies monitoring of AI agents, applications, and infrastructure in a single solution.

Missing alt text value

One experience for everything you run

Your applications and AI agents finally share one view. CloudWatch Omni converges infrastructure and agent observability into a single experience, delivered in your IDE or Omni web UI with enterprise SSO. Auto-discovered topology maps services, dependencies, and golden metrics across accounts and regions without manual configuration. Instrument once via OTLP or bring existing CloudWatch telemetry with zero reconfiguration. No AWS Console required for end users. 

Missing alt text value

Ship agents with confidence

Eval-driven observability gives you continuous quality scores on every agent invocation, from correctness and groundedness to tool selection accuracy. Trace full execution paths and catch silent regressions before your customers do. Run online evaluations against production traffic and offline evaluations against versioned datasets capturing edge cases. Quality scores live as first-class signals alongside your traces and metrics, not in a separate tool. Use Amazon Bedrock Agentcore built-in evaluators or bring your own. 

Missing alt text value

Bring your own stack

Instrument once with OTLP and keep the agent frameworks you already use, like LangChain, CrewAI, Strands, and OpenAI SDK. Choose from AgentCore built-in evaluators, Braintrust, DeepEval, Ragas, or custom LLM-as-a-judge. Use the same evaluator interface in development and production so your investment carries forward. Python and TypeScript are supported natively, with any runtime reachable via OTLP and no proprietary formats or vendor lock-in.

Missing alt text value

Answers from the start

Stop building dashboards just to ask a question. Application topology shows agents alongside the services, APIs, and infrastructure they depend on. Ask in plain English and let the AI assistant guide your investigation from symptom to root cause. Get SQL and PromQL generated for you across metrics, traces, and logs. Native alerting fires on conditions across your telemetry and routes to Amazon SNS and Slack.

Missing alt text value

Build with AI from the ground up

Every observability task has intelligence behind it. CloudWatch Omni generates queries, correlates signals across metrics, traces, and logs, and drives investigations powered by AWS DevOps Agent. Engineers and the AI assistant investigate together in shared sessions, so the next person joins with full context already in place. The system automatically discovers services, maps dependencies, and adjusts alarms as your applications evolve.

Missing alt text value

Customers and design partners

Sony Group Corporation

“At Sony, our enterprise-wide agentic AI platform is now supporting hundreds of proof-of-concept and production workloads. At this scale, observability and evaluation are essential. With Amazon CloudWatch Omni, I can go from a single trace straight to evaluation, AI analysis, comparison, or dataset creation, it's all right there. The ability to create a dataset from live traces with a single click was especially impressive. Assembling datasets for evaluations is often a bottleneck on the business side, and being able to create them directly from traces significantly lowers that barrier. Additionally, being able to work within my development environment seamlessly across local and cloud data made the whole experience very easy to use.”

Masahiro Oba, Senior General Manager of AI Acceleration Division, Digital & Technology Platform, Sony Group Corporation 

Crest Data

“Amazon CloudWatch Omni provides strong end-to-end observability for AI agents that our customers require. Zero-config trace capture is reliable and includes multi-agent support, no manual setup, and IDE support. Token counts aggregate accurately across instrumentation conventions, and the Evaluate pillar, which spans dataset curation, LLM-as-judge evaluators, and human annotation, works end to end. Overall, Amazon CloudWatch offers a robust foundation with comprehensive observability capabilities and significant potential for continued innovation in the AI agent observability space.”  

Nirav Kelaiya, Director of Engineering, Crest Data 

RESTRICTIONS Crest Data Logo

Capital One

"Capital One operates one of the largest observability footprints in financial services. As a design partner for Amazon CloudWatch Omni, we helped shape a single AI-powered observability solution that will give our engineers topology-aware intelligence and natural-language querying across all telemetry from a single surface, with full data ownership through OpenTelemetry. It's a fundamentally different operating model for AIOps and resilience at scale."   

Parvez Naqvi, Managing Vice President, Cloud Platform & Resilience Engineering, Capital One 

The Capital One logo featuring the company name with a red swoosh design above the text.

Use cases

See everything you run in one view

Auto-discovered topology maps services, dependencies, and golden metrics across accounts and regions. Ask questions in natural language and let AI guide your investigation, off AWS Console with enterprise SSO. CloudWatch Omni organizes dashboards around services and user journeys rather than individual resources, giving you an application-centric view that reflects how your teams actually think about their systems. 

Whether your workloads span a single account or hundreds across multiple regions, topology is built automatically with no manual configuration or tagging required. Operators get a unified starting point that eliminates the need to mentally stitch together signals from disconnected consoles. 

Explore CloudWatch Omni

Query all your telemetry with AI-powered investigation

Query logs, metrics, and traces with SQL, PromQL, or plain English. The AI assistant correlates signals and guides you from symptom to root cause. AWS DevOps Agent acts as your built-in investigator. Ask it a question like “Why did latency spike on the checkout service?” and it pulls the relevant logs, metrics, and traces together, walking you through the analysis step by step. 

Native alerting routes findings to Amazon Simple Notification Service and Slack so your team is notified the moment something degrades. Instead of context-switching between separate tools for each telemetry type, you search across all signals in one place and let AI surface the connections you would otherwise spend hours finding manually. 

Explore Omni agent

Ship faster by evaluating against real-world traces, not synthetic tests

Evaluations run continuously against live traffic, so you catch quality drops in alerts before your customers do. Operators get fleet-wide visibility into latency, errors, cost, and eval scores as first-class signals. 

Run online evaluations against production traffic and offline evaluations against versioned datasets that capture edge cases, powered by Amazon Bedrock AgentCore along with integrations like Braintrust, DeepEval, Ragas, and custom LLM-as-judge evaluators. Measure correctness, groundedness, hallucination, toxicity, and tool selection accuracy on every agent invocation. Quality scores sit alongside your traces and metrics as first-class observability signals, not in a separate tool, so you can set alerts on eval regressions the same way you alert on error rates or p99 latency. 

Explore agent observability

Fix broken agents in seconds with full execution context

When something fails, get traces, session paths, and root cause right where you code. AI-assisted debugging pulls from real production signals, so you resolve issues faster without context-switching. CloudWatch Omni delivers extensions for VS Code, Kiro, and Cursor that bring full observability into your development workflow. 

Trace complete agent execution paths, from every LLM call, tool invocation, to reasoning step, directly in your editor. With support for Python and TypeScript across frameworks including LangChain, LangGraph, CrewAI, OpenAI Agents SDK, Vercel AI SDK, and Strands, you can instrument, test, and debug agents without leaving the environment where you build them. 

Explore the IDE extension

Observe agents and the applications they interact with together

Agents call APIs, retrieve from knowledge bases, and run on infrastructure. CloudWatch Omni shows agents alongside application components in a single view with auto-discovered topology and golden metrics. Full-stack correlation connects agent reasoning to application health to infrastructure, so when an agent produces a poor response, you can trace the issue from the eval score through the retrieval call, into the downstream API, and down to the compute layer — all without switching tools. 

This unified view is critical because agents don’t operate in isolation; they depend on the same services, databases, and networks your applications do, and failures in any layer can cascade into degraded agent behavior that traditional siloed monitoring would never surface. 

Explore and query telemetry

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages