AWS for Industries
Building self-learning agents for intelligent Network Operations with Kiro Crew
A self-learning agent does more than remember a factual answer. It persists reusable knowledge, extracts knowledge when needed, applies it to future work and produces reusable operating procedures for automation.
Communication Service Providers’ (CSPs) network operations knowledge lives across data and people. Network Telemetry, alarms, topology, configuration, and incident records all comprise the Telco Data Lake. Given the complexity across various system silos, there is a unique perspective network Engineers (humans) contribute with their experience to add runtime context, steer the agent in the right direction, and inject knowledge which aids in distinguishing symptoms from causes when existing data sources are not enough to determine a root cause or make an informed decision. A self-learning system must preserve context from human-to-agent and agent-to-agent interactions and apply the combined learning to future investigations.
To build such a self-evolving loop for Autonomous Network Operations, agentic systems must react to network events, drive an investigation loop, remember context, learn how its operator works, and coordinate across unique tools and workflows, so the system comes back with progress instead of another workflow to restart. This blog will focus on the persistent and self-learning aspects of the Agent system.
Persistent and self-learning agents
Persistent context is designed to carry forward diagnostic methods so later investigations can start from prior reasoning rather than from scratch. Say an event in the system triggers an agent loop, and the agent analysis results in competing hypotheses, an engineer can provide a decision test. The agent collects the additional evidence and either reaches a supported conclusion or preserves the uncertainty. The system can retain the test with its provenance, confidence, scope, and review date. After the procedure is successfully demonstrated in a later investigation, the system can stage the procedure as a candidate skill for human review.
Figure A: A high-level agent loop with persistent local memory
The diagram shows an event-driven system that processes network events through an agent loop. The agent evaluates competing paths, reaches a decision, and stores the outcome in persistent memory (semantic, learning, and history). Successful procedures can be crystallized into reusable skills in the Skill Library.
Kiro Crew is one such implementation of an open-source persistent AI agent workspace designed to run complex, multi-step tasks. Kiro Crew runs on Kiro CLI via the Agent Client Protocol. Each investigation session is a Kiro CLI agent session with Kiro Crew supplying the tools, local memory, and overall orchestration around the system. Existing Kiro configuration, including steering files, skills, and custom agents, carries over without extra setup. Kiro Crew has persistent memory that survives across sessions. It remembers the user preferences, project context, daily activity, and corrections you teach it.
This post focuses on two Kiro Crew self-learning capabilities: durable lessons that preserve explicit and implicit corrections steered by the human Network Engineer and dynamic skill creation that stages reusable operating procedure (skill) for review. Kiro Crew interacts with a Network Engineer in their daily messaging channel – could be Slack, Webex or other existing channel. Additional Kiro Crew capabilities fall outside the scope of this illustrative example. We focus on a learning loop for a sample cable access network issue diagnostic scenario. The scenario is illustrative; it is designed to show how the learning loop works. The investigation agent performs the diagnosis. Kiro Crew starts the work, supplies approved tools, delivers the result, records explicit human feedback, and brings that feedback into later sessions. The analysis of Kiro Crew’s lesson and skills come from its open-source implementation. A Communication Service Provider could use this reference for a future implementation customized to operators’ needs and workflows.
Let us go through an example scenario of a Kiro Crew Agent powered investigation, steered by a Network Engineer, and the system self evolves via a learning loop.
An overnight investigation
Figure B: Sample self-learning network investigation loop with Kiro Crew
The diagram traces three stages of maturity: Night 1, where the agent investigates fiber node FN-114 with engineer guidance; Night 2, where it applies lessons learned to a new incident on FN-208 and stages a candidate skill; and Night N, where approved skills enable autonomous response with minimal oversight.
Let us trace this scenario in detail.
Night 1: At 02:14, an existing monitoring system detects upstream signal-quality degradation event on a certain fiber node (FN) (1). Kiro Crew starts an investigation session (2), spawning a network investigation agent while nobody is on shift (3).
Kiro Crew coordinates the investigation session, connects the agent to operational data through approved Model Context Protocol (MCP) tools, and supplies lessons from network engineers to future sessions.
Through approved MCP tools, the agent examines operational telemetry, correlated access-network alarms, topology, and recent configuration history. The data shows a cluster of modems with declining upstream quality. Headend systems show no corresponding fault. The topology graph maps every affected modem to the same amplifier cascade, and configuration history shows no relevant change.
The investigation agent posts its assessment to the operations messaging channel for example Slack, Webex or similar:
Figure C: Night 1 agent assessment, identifying two competing hypotheses for FN-114 upstream impairment
The agent’s posted assessment narrows the affected modems to a shared amplifier (AMP-114-03) but acknowledges that the available evidence does not distinguish between noise ingress and return-path equipment degradation. The agent escalates for human review.
The topology graph did useful work. It reduced a group of affected modems to one shared element. The graph shows where the impairment spread, not where it entered, the real root cause.
At 07:40, a network engineer reads the digest and replies with a technical guidance (4):
Figure D: Engineer steering, a diagnostic test to distinguish the two hypotheses
The engineer’s correction teaches the agent a diagnostic method: upstream-only degradation with normal downstream quality points to ingress or return-path faults, while bidirectional degradation indicates a shared-infrastructure problem. The instruction is to sectionalize the return path before recommending amplifier replacement.
In hybrid fiber-coaxial networks, sectionalizing the return path means isolating and testing progressively smaller branches of the upstream plant while watching whether the impairment disappears. This process helps distinguish a faulty amplifier from ingress introduced by a loose or corroded connector, cracked cable, unterminated port, damaged drop, or another return-path-specific fault.
In this illustrative scenario, the downstream measurement was already available. The agent did not need another data source; it needed to know which comparison was diagnostic. The rule is explicitly local, injected via human steering. It reflects one plant, one amplifier vintage, and that engineering team’s knowledge history. Because the engineer gave an authenticated, steering instruction, the agent records it as a durable Kiro Crew lesson. The lesson persisted through Kiro Crew lessons. The lesson could be explicit where the user says “Remember this” or implicit, which is extracted from the session during history consolidation. The lesson becomes available to later sessions without retraining the model or creating a labeled dataset.
The next night (Night 2 in the diagram above)
At 01:50 the following night, monitoring detects the same initial upstream degradation signature on FN-208 (5). This time, Kiro Crew supplies the relevant lesson to the new investigation. The agent first verifies its applicability – the agent verifies that FN-208 uses the same plant design, return-path profile, and amplifier family as FN-114. It then compares upstream and downstream measurements over the same interval. This is done through Kiro Crew’s native memory that is incorporated into the Agent working context.
The investigation Agent applies the learning (6), checks downstream quality early. This time, upstream and downstream quality degrade during the same window. The agent posts:
Figure E: Night 2 agent assessment, applying the learned diagnostic test to FN-208
On the second night, the agent applies the learned test to FN-208 and finds bidirectional degradation, pointing to a shared-plant or active-equipment fault rather than simple ingress. It recommends human review and a dispatch to test the amplifier, power supply, connectors, and coax cable.
In this scenario, the same test leads to a different conclusion because the evidence is different. The lesson did not tell the agent what the fault was. It told the agent what to compare, applying the lesson it had learned from the human steering in the past.
A stored answer risks applying yesterday’s conclusion to today’s incident. A stored diagnostic test gives the next investigation a method for choosing among hypotheses.
After the second investigation, Kiro Crew identifies the demonstrated sequence as a candidate repeatable skill. With dynamic skill creation enabled, Kiro Crew evaluates the completed session and stages an identified procedure as a candidate skill. Under the default approval setting, the candidate skill remains inactive until a human reviews it. A lesson preserves the engineer’s immediate correction. A skill turns the demonstrated sequence into an inspectable operating procedure that the Network Engineer can choose to crystallize (7). Steps (8) and (9) in figure B represent future reuse of lessons and approved skills automating these repeatable procedures.
What Kiro Crew adds
The monitoring system detects the anomaly, and the investigation agent performs the diagnosis. Kiro Crew completes the operational loop.
Unattended orchestration. Event driven, authenticated webhooks and schedules start investigations before an engineer reviews an incident. Kiro Crew preserves session context, delivers a digest through an existing messaging channel, and starts the next investigation with relevant prior knowledge. The engineer reviews prepared evidence instead of reconstructing the incident from raw alarms.
Human feedback across sessions. The correction arrives where the agent posted its findings. One explicit instruction becomes a durable lesson, and a later session applies it to different evidence. The field organization needs no separate labeling application.
Controls and ownership. Dynamic skill creation is off by default, and every script-bearing candidate requires human approval. Additional controls (sandbox, denied-by-default commands, credential redaction, audit log) are documented in the Kiro Crew memory and skill design.
Inspectable knowledge. Lessons and generated skills remain visible instead of disappearing into model weights. Operators inspect, revise, remove, and export accumulated knowledge. Because Kiro Crew is open source, teams also inspect and extend the mechanisms that assemble agent context.
Kiro Crew runs locally on your own hardware or remotely on servers (such as an EC2 instance or Docker container). You can work with it from the desktop app, web dashboard, and CLI, or continue the same work through connection tools like Slack and other messaging applications.
Kiro Crew’s differentiation is not basic memory storage or retrieval; equivalent persistence and retrieval can be provided by an external memory service. Its value is the integrated operating loop: capturing explicit human corrections as inspectable lessons, carrying relevant context across sessions, coordinating agents and tools, and converting demonstrated procedures into skill candidates that remain pending for human review by default. External domain memory can complement Kiro Crew through approved tools or underpin a separately implemented loop- when operational knowledge requires structured metadata, domain-aware scoping, and application-managed lifecycle controls.
Design persistent knowledge with scope and lifecycle controls
A diagnostic test that helps resolve one incident should not be applied blindly to every scenario. Its validity may depend on the market, plant design, vendor, hardware generation, firmware version, or time period. Operational teams should therefore record the conditions under which guidance applies and establish when an engineer must review it. Kiro Crew can preserve these conditions as part of a lesson, while the investigation procedure verifies them against the current network evidence before applying the guidance.
Kiro Crew lessons distinguish standing instructions from topic-specific findings and allow engineers to correct, replace, or remove stored guidance. Topic-specific findings can be withheld when the current request does not match them, while standing lessons remain eligible regardless of topic. Lessons do not expire automatically or carry a native review date, so operator review remains the safeguard against outdated guidance.
Alongside lessons, Kiro Crew uses episodic memory to retrieve relevant fragments from previous events and conversations. Relevance filtering, recency-based scoring, and diversity reranking help surface useful context without replaying entire sessions. Episodic memory helps the agent recall what happened, while lessons preserve operator corrections and reusable guidance. The decay mechanisms applied to episodic memory do not expire explicit or implicit lessons. Separately, Kiro Crew allows the autonomous, persistent AI agent to self-evolve by synthesizing recurring user patterns into reusable, named Markdown files (SKILL.md). Skills auto-generated still require human review and approval, before they are added to the library of available skills for the Agent.
Conclusion
The value proposition of a self-learning operations system is that an engineer’s correction can change how future work is performed. In the first investigation, the agent preserved uncertainty because two hypotheses fit the available evidence. The engineer then supplied the missing diagnostic test: compare upstream and downstream quality over the same interval. In a later investigation, the agent verified that the lesson’s conditions applied, used the same test with different evidence, and produced an assessment consistent with the new evidence. The system improved not by memorizing an earlier conclusion, but by preserving and reusing a method.
Kiro Crew does not replace monitoring systems, investigation agents, or network-domain knowledge systems. It coordinates the operational loop across tools, sessions, human feedback, and inspectable knowledge reuse. The pattern extends beyond cable networks: teach the agent a conditional diagnostic test, state the conditions under which it applies, verify those conditions before reuse, and promote demonstrated procedures to reusable skills only through human review.
Call to Action
Start with a real operational question: have Kiro Crew start with a small problem, say investigate logs, correlate the available evidence, and return a supported assessment. As the engineer reviews the result, they can add context, correct an assumption, or teach the agent which diagnostic test to apply. Kiro Crew preserves that guidance across sessions, allowing the user to observe how persistent context carries forward into later work and how a useful investigation sequence can become a governed, reusable skill. The experience begins with solving today’s problem and naturally reveals how the system learns to handle tomorrow’s work more effectively. Kiro Crew is open source, so you can extend it to fit your tools, workflows, and preferences. Operators can configure their own agents, connect domain-specific MCP tools, shape the learning loop around their operational environment and benchmark results.




