AWS Cloud Operations Blog
Best practices for writing AWS DevOps Agent Skills
When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the system. The result: investigation quality varies depending on who gets paged. When the engineer who built the service is available, root cause analysis takes significantly less time. When they’re not, the same investigation can take hours or miss the root cause entirely.
AWS DevOps Agent Skills address this gap. They let you encode your team’s best investigation workflows into modular instruction sets that the agent loads automatically during incidents. Instead of relying on whoever happens to be on call knowing the right steps, Skills make that expertise available to the agent 24/7. Consistent, repeatable, and independent of who is paged.
A well-written skill turns tribal knowledge into a reusable investigation playbook. A poorly written one gets ignored entirely. In this post, we share best practices for building Skills that reliably activate, follow structured investigation workflows, and compose with other skills to reduce mean time to resolution (MTTR).
One global financial services firm we worked with saw the shift firsthand. After encoding their team’s investigation shortcuts into Skills, the agent went from “interesting demo” to first responder on their P1 incidents. Surfacing likely root causes that would otherwise take a senior engineer significant manual log correlation to find. Their takeaway matched ours: the quality of your Skills determines the quality of your investigations.
Background: what are Skills?
Skills are self-contained directories with a required SKILL.md file and optional reference materials such as architecture diagrams, metric threshold tables, or troubleshooting flowcharts. They follow a subset of the open Agent Skills specification, supporting non-executable documents: Markdown, PDFs, images, and data files.
Every skill needs frontmatter (a metadata block at the top of SKILL.md) containing a name and description. The agent evaluates the description to decide whether the skill is relevant to the current task. The rest of the file contains the investigation instructions themselves.
my-skill/
├── SKILL.md # Required: instructions + frontmatter
├── references/ # Optional: metric tables, error code mappings
└── assets/ # Optional: diagrams, flowcharts, data files
If you’ve already deployed Agent Spaces following the guidance in Best practices for deploying AWS DevOps Agent in production, Skills are the next step as they specialize the agent’s behavior within those spaces.
Without Skills, the agent investigates using its general knowledge of AWS services. It can still query metrics, review logs, and check deployments. But it doesn’t know your team’s investigation shortcuts, your environment-specific thresholds, or which metrics matter most for your services. Skills close that gap. They give the agent a head start by pointing it directly at the checks that matter, in the order that matters, skipping the trial-and-error that slows down a general-purpose investigation.
When to use a skill versus an agent instruction
Skills are not the only way to shape how AWS DevOps Agent behaves. Agent instructions, stored as an AGENTS.md file and configured on the Knowledge page of your Operator Web App are the other lever, and the two solve different problems. Getting the choice right keeps the agent focused and protects its working memory.
The difference is when the guidance loads. Agent instructions are always on: the agent service injects them into the system prompt at the start of every session, no matter what the agent is working on. Skills load on demand. The agent reads a skill only when its description matches the task in front of it. So agent instructions define how the agent behaves in every session, while a skill gives it a specialized playbook for a specific situation.
That single difference drives a simple rule. If the guidance must apply to every investigation, for example, response formatting, a security policy like never surfacing secrets in plaintext, or a house rule to always check recent deployments before proposing a root cause. Then it belongs in agent instructions, where you can guarantee it is present. If the guidance is a procedure for one scenario, such as how to investigate RDS connection exhaustion, or an ECS crash loop. It belongs in a skill, where it loads only when relevant and stays out of the context window the rest of the time.
The practical differences reinforce the split:
Agent instructions (AGENTS.md) |
Skills | |
|---|---|---|
| When it loads | Every session, unconditionally | On demand, when the description matches the task |
| Best for | Standing policy — formatting, security, always-do checks | Scenario procedures — investigation playbooks |
| How it’s selected | You scope it to all agents or one agent type | The agent matches your description at runtime |
| How many | One per agent (global, or per managed agent) | Many per Agent Space |
| Extra files | Markdown only, no attachments | Markdown plus reference files, images, and data |
| Size discipline | Hard limit 25 KB; keep it lean (~120 lines recommended) | Loads only when needed, so keep each one focused |
A quick test: if you’d want the guidance in front of the agent even before it knows what it’s investigating, that’s an instruction. If it only makes sense once you know the incident is “an RDS latency problem,” that’s a skill.
One caution on size. Because instructions load on every session, they compete with your questions, the logs the agent reads, and its own reasoning for a fixed amount of working memory. The service guidance is to keep them short and move specialized procedures into skills. A sprawling AGENTS.md that tries to encode every investigation procedure is the anti-pattern. That content should be split into targeted skills.
Skills and instructions are two of several ways to extend the agent. Custom agents bundle a system prompt, tools, and skills into a purpose-built workflow such as a scheduled health report, and MCP servers add custom diagnostic tools. The open-source AWS DevOps Agent Tools repository has ready-to-use examples of all three: skills, custom agents, and MCP servers that you can import as-is or adapt as a starting point for your own.
Write descriptions that activate reliably
The description field in your frontmatter determines whether the agent loads your skill. A vague description means the agent skips the skill entirely, even if the instructions inside are excellent. This is the single highest-impact area to get right.
Write the description from the agent’s perspective. Include the specific services, error types, symptoms, and scenarios that should trigger activation.
Example — too vague:
---
name: database-skill
description: Helps with database issues.
---
Example — specific and actionable:
---
name: rds-connection-exhaustion
description: Investigation procedures for Amazon RDS and Amazon Aurora
connection exhaustion, including max_connections limits, connection pool
misconfiguration, and idle connection accumulation. Use this skill when
investigating DatabaseConnections alarms, connection timeout errors, or
"too many connections" application errors.
---
Think about what an on-call engineer would search for during an incident. Those terms like the alarm names, error messages, and symptoms belong in your description.
A good test: read the description and ask yourself, “If I were triaging an incident, would I know from this description alone whether this skill applies?” If the answer is no, add more specifics.
We saw this firsthand. We had a skill called db-health with the description “Monitors database health.” When a production RDS instance hit max connections at 3 AM, the agent didn’t load it because nothing in “monitors database health” matched the alarm RDS-DatabaseConnections-Critical firing in the incident. We updated the description to: “Investigation procedures for RDS connection exhaustion, max_connections breaches, and DatabaseConnections alarm spikes.” Same skill, same instructions inside. But now the agent loads it every time that alarm fires.
Structure instructions as investigation steps
Skills should read like a senior engineer’s investigation playbook, not a list of facts. Each step should tell the agent what to check, what to look for, and what to do next based on findings.
The difference between a skill that works and one that doesn’t often comes down to whether you’ve included decision logic. Here’s an example that fails, followed by one that succeeds:
A skill that fails — vague, passive, no decision logic:
---
name: ecs-skill
description: Helps investigate ECS issues.
---
# ECS Investigation
Check ECS task status. Look at CPU and memory metrics.
Review recent deployments. Check the load balancer.
If something looks wrong, investigate further.
This skill has two problems. The description is too broad to activate reliably. “ECS issues” matches everything and nothing. The instructions are a checklist of nouns, not an investigation workflow. There’s no sequencing, no thresholds, no decision points telling the agent what “wrong” means or what to do next.
A skill that works has specific activation, structured logic, and clear decision routing:
---
name: ecs-task-crash-loop
description: Investigation procedures for ECS tasks entering crash loops
(repeated STOPPED status with non-zero exit codes), including OOM kills,
health check failures, and container dependency issues. Use when
ECS-TaskFailure alarms fire or tasks cycle between PENDING and STOPPED.
---
# Step 1: Identify the failure pattern
Query ECS stopped tasks for the affected service over the past 2 hours.
Categorize by stop reason:
- Exit code 137 → OOM kill. Proceed to Step 2.
- Exit code 1 with "essential container exited" → dependency failure.
Proceed to Step 3.
- No exit code + "health check failed" → container healthy but
unresponsive. Proceed to Step 4.
# Step 2: Investigate memory exhaustion
Pull CloudWatch Container Insights memory utilization for the task
definition. Compare against the task's hard memory limit. If utilization
consistently exceeds 90% of the limit in the 30 minutes before failure,
recommend increasing the memory allocation and flag the most recent
deployment for memory-affecting changes.
# Step 3: Investigate dependency failures
Identify which container exited first by comparing task state change
timestamps. Check that container's logs for startup errors — common
causes include missing environment variables, failed health checks on
dependent services, or image pull failures from a registry outage.
# Step 4: Investigate health check timeouts
Compare the health check configuration (interval, timeout, retries)
against the container's actual startup time. If the container needs
longer to initialize than the health check allows, it will be killed
before it becomes healthy. Check if a recent deployment increased
startup time (new dependencies, larger models to load, database
migrations on startup).
Notice the pattern: each step has a clear action, specific things to look for, and a decision point that routes to the next step. The agent follows this like a flowchart, not a reading list.
Avoid writing skills that only list facts or metrics without telling the agent what to do with them. A table of Amazon CloudWatch metric thresholds is useful as a reference file, but the SKILL.md itself should contain the investigation logic that uses those thresholds.
Target the right agent types
Not every skill applies to every phase of an investigation. AWS DevOps Agent supports multiple agent types, and you can target skills to specific ones using the Agent Type setting when creating or uploading a skill.
Generic is the default — it makes the skill available to all agent types and is the right choice when starting out. The other agent types correspond to distinct phases of the investigation and delivery lifecycle:
| Agent type | When it runs | What it does | Example skill |
|---|---|---|---|
| Incident Triage | Immediately on incident arrival | Correlates new incidents with active investigations, classifies severity, decides whether to investigate or skip | Scheduled maintenance skip rules; severity classification based on affected service tier |
| Incident RCA | After triage decides to investigate | Deep investigation — queries metrics, logs, traces, and deployments to identify root cause | RDS connection exhaustion playbook; ECS crash-loop analysis |
| Incident Mitigation | After root cause is identified | Generates step-by-step remediation plans with immediate fixes and long-term prevention | Deployment rollback procedures; connection pool scaling guidance |
| On-demand | When explicitly invoked via chat | Responds to ad-hoc questions and tasks outside the incident lifecycle | Architecture documentation queries; capacity planning calculations |
| Evaluation | Runs periodically or on schedule | Proactive assessments of infrastructure health, configuration drift, and observability gaps | Monthly observability gap analysis; security posture checks |
Each skill consumes context when loaded. If the agent loads a detailed mitigation playbook during initial triage, that’s context spent on instructions it doesn’t need yet. Targeting reduces this overhead and keeps the agent focused on the current phase.
A practical approach: start with Generic for all your skills. After running several investigations, review which phases each skill was most useful in. Then narrow the targeting. A deployment rollback procedure belongs in Incident Mitigation. Loading it during triage wastes context on steps the agent can’t act on yet. An observability gap analysis belongs in Evaluation. This is proactive work, not incident response. A severity classification guide belongs in Incident Triage. It needs to run immediately when an alarm fires, not after the agent has already spent twenty minutes investigating.
The distinction between On-demand and Generic is important: On-demand skills only activate when someone explicitly asks the agent a question through chat. A skill documenting your team’s architecture for ad-hoc queries belongs in On-demand. A skill that should fire automatically during incidents belongs in one of the incident types or Generic.
Mitigation skills are worth revisiting now that AWS DevOps Agent supports directed actions. Directed actions are operations you explicitly ask the agent to perform against connected services and AWS accounts. Read-only actions are available by default; actions that create or modify resources are disabled by default and must be enabled deliberately, with every approval and action attributable to the approving operator in AWS CloudTrail. If you enable directed actions, an Incident Mitigation skill can move from describing a remediation to walking the agent through executing an approved one. Write those skills with the same discipline as the rest: explicit preconditions, a clear stopping point, and instructions to surface the plan for approval before it acts.
After narrowing our skills to specific agent types, we observed that the agent’s investigation outputs became more focused. Triage responses were concise severity assessments rather than premature root cause speculation, and RCA phases went deeper because the context window wasn’t consumed by mitigation steps the agent didn’t need yet.
Include reference materials
Skills with supplementary reference files give the agent structured data to reason over during investigations. The SKILL.md contains the investigation logic; reference files contain the domain knowledge that logic operates on.
Metric threshold tables define what “normal” and “concerning” look like for your environment. The agent can compare live metrics against these baselines during an investigation.
# references/rds-metrics-reference.md
| Metric | Normal range | Investigation threshold |
|---------------------|------------------------|--------------------------|
| DatabaseConnections | < 70% max_connections | > 80% max_connections |
| ReadLatency | < 5ms | > 20ms |
| WriteLatency | < 5ms | > 20ms |
| FreeStorageSpace | > 30% total storage | < 20% total storage |
| CPUUtilization | < 70% | > 85% |
Error code mappings translate cryptic codes into investigation paths. Instead of the agent guessing what ORA-12519 means, a reference file can map it directly to “listener connection limit reached: check connection pool configuration.”
Architecture context documents service dependencies, data flows, and failure domains specific to your environment. This is especially valuable for microservice architectures where the blast radius of a failure isn’t obvious from the infrastructure alone.
Escalation procedures tell the agent when and how to recommend engaging other teams. This is operational knowledge that rarely exists in formal documentation but is critical during incidents.
The directory structure for a skill with reference materials looks like this:
ecs-deployment-investigation/
├── SKILL.md
├── references/
│ ├── ecs-error-codes.md
│ ├── deployment-strategies.md
│ └── healthy-thresholds.md
└── assets/
└── ecs-investigation-flow.png
Compose skills for end-to-end workflows
Individual skills are useful on their own, but the real value comes from composition. AWS DevOps Agent reads multiple skills during a single investigation, so you can design skills that complement each other without duplicating content.
Consider an Amazon ECS service that starts failing after a deployment. Three separate skills can work together: one that retrieves recent deployments from your CI/CD system, one that searches code repositories for relevant changes, and one that investigates ECS task failures and scaling behavior. The agent loads all three and correlates deployment timing with code changes and service health, the same workflow a senior engineer would follow, but automated.
The design principle is straightforward: each skill should be independently useful but composable with others. A CI/CD pipeline skill shouldn’t assume the code repository skill is also loaded. It should produce findings that are useful on their own and richer when combined with findings from other skills.
This also means you should avoid building monolithic skills that try to cover an entire investigation end-to-end. A single skill that covers “everything about ECS” becomes too broad for the agent to activate reliably (the description can’t be specific enough) and too large to load efficiently. Split it into focused skills such as ECS deployment issues, ECS scaling issues, ECS networking issues and let the agent compose them as needed.
Guide the agent in using custom MCP tools
If you’ve connected custom MCP servers to AWS DevOps Agent, skills can document how to use those tools effectively. Without a skill, the agent knows the tool exists but may not know the right parameters for your environment or how to interpret the results.
---
name: custom-deployment-tracker-investigation
description: Procedures for investigating incidents related to recent
deployments using the internal deployment tracking MCP integration.
Use when incidents correlate with service deployments or configuration
changes pushed through the CI/CD pipeline.
---
# Deployment correlation investigation
When investigating incidents that may relate to recent deployments:
## Step 1: Query recent deployments
Use the `deployment-tracker-get-changes` tool with these parameters:
- `environment`: Match the affected environment (production, staging)
- `timeRange`: Set to 4 hours before incident start
- `service`: Use the affected service identifier
## Step 2: Correlate deployments with the incident timeline
Compare deployment timestamps against the incident start time.
Deployments within 30 minutes of incident onset are high-priority
suspects. Check whether the deployment included configuration changes
or dependency updates — these are common causes of latency-related
incidents that appear after traffic ramps post-deployment.
This pattern documents when to invoke the tool, what parameters to use for different scenarios, and how to interpret the results. The agent goes from “I have access to a deployment tracker” to “I know how to use this tool to investigate deployment-related incidents in this environment.”
Skills and memory work together
Skills are the instructions you author; memory is the operational knowledge the agent accumulates and that you can curate. AWS DevOps Agent now maintains learned knowledge such as your environment topology, code dependencies, pipeline structure, and tool-use patterns as memory, and you can create your own memory stores to group operational knowledge for a team, a service, or a recurring problem. The two are complementary: a skill encodes the investigation workflow, while a memory store holds the environment-specific facts that workflow draws on.
A practical division: put the durable, reusable procedure in a skill, and let fast-changing environment facts live in memory where they can be updated without editing the skill. This keeps skills stable and reduces the “skill rot” problem described in the next section, because the details most likely to drift are the ones you’ve moved into memory.
Failure modes that aren’t obvious
The basics such as writing specific descriptions, using decision trees, don’t build monolithic skills are covered in the sections above. The failures below are subtler. They emerge after you’ve been running Skills in production for weeks or months.
Skills that conflict when composed. Two skills can give contradictory guidance when loaded together. A “scale up database connections” skill and a “reduce connection pool size” skill might both activate during the same RDS incident, leaving the agent with conflicting recommendations. Before publishing a new skill, review what other skills target the same alarms or symptoms. If two skills might load simultaneously, make their decision boundaries explicit: “If connection count is high but CPU is normal, this is a pool leak — reduce connections. If connection count is high and CPU is saturated, the database needs more capacity.”
Skills that rot as infrastructure evolves. A skill written for your ECS cluster six months ago references task definitions, service names, and metric thresholds that may no longer exist. Unlike code, skills don’t fail loudly when they reference stale infrastructure. The agent simply follows outdated steps and produces irrelevant findings. Treat skills like runbooks: review them quarterly, tie them to your change management process, and add a last_verified comment in the SKILL.md so reviewers know when the investigation steps were last tested against live infrastructure. For the details most prone to drift like thresholds, resource names, dependency maps, consider moving them into a memory store so they can be updated independently of the skill.
Overly rigid steps that prevent the agent from adapting. A skill that prescribes “check metric X, then check metric Y, then check metric Z” in strict sequence can trap the agent in a linear path when the actual incident doesn’t match the expected pattern. If Step 1 returns normal results, the agent should have a branch that says “skip to Step 4” or “this skill may not apply. Report findings so far and defer to general investigation.” Without escape routes, the agent dutifully completes all steps and reports a clean bill of health while the actual root cause sits in a system your skill never mentions.
Skills that duplicate each other’s scope without realizing it. Over time, different team members write skills that overlap like “API latency investigation” skill and a “service timeout investigation” skill that both check the same CloudWatch metrics and trace the same call paths. The agent loads both and performs redundant work, consuming context on duplicate findings. Maintain a skill inventory (even a simple spreadsheet mapping skills to the alarms and services they cover) and audit for overlap when adding new skills.
Getting started
Start with your top three incident categories from the past quarter. These are the scenarios where a skill will have the most immediate impact.
For each one, talk to the senior engineer who usually handles that type of incident. Ask them: “When you get paged for this, what’s the first thing you check?” That answer becomes Step 1 of your skill. “What do you check next? What tells you it’s X versus Y?”, and you have your investigation workflow.
A 20-line SKILL.md with clear, specific steps is more effective than a 200-line document with vague guidance. Start simple, run investigations, and refine based on what the agent does well and where it gets stuck. As your skill library grows, you can manage skills and other assets declaratively as infrastructure as code and deploy them across Agent Spaces through a pipeline, the same way you manage the rest of your infrastructure.
If you’d rather not start from a blank page, the open-source AWS DevOps Agent Tools repository has community skills, custom agents, and MCP servers that you can import directly or adapt.
For more details on skill structure and creation, see the AWS DevOps Agent Skills documentation. For guidance on setting up Agent Spaces, see Best practices for deploying AWS DevOps Agent in production.
Skills are how you stop losing institutional knowledge when engineers change teams. Every incident your team investigates well is a skill waiting to be written. Start with the investigation your best engineer runs on autopilot. The one where they already know which dashboard to open, which log group to query, and which metric crossing which threshold means “rollback now.” That’s your first skill. Once it’s written, that expertise runs at 2 AM regardless of who gets paged.