Get in Touch

Course Outline

Introduction to Agentic AI in Operations

  • Transitioning from static runbooks to reasoning agents: the evolution of IT automation
  • Agent components: reasoning loops, tool utilisation, memory, and planning
  • Determining when to automate versus when to retain human involvement

Agent Frameworks and Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
  • Multi-agent architectures: supervisor, hierarchical, and swarm models
  • Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent solutions
  • Creating your first operational agent: querying monitoring, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
  • Agent-based log querying: integration with Elasticsearch, Loki, and Splunk
  • Utilising infrastructure tools: kubectl, Terraform, and Ansible via agent actions
  • Designing secure tool interfaces with parameter validation and idempotency

Incident Response Automation

  • Automating incident triage: severity classification and routing
  • Generating root cause hypotheses and collecting evidence
  • Automated remediation: executing restart, scaling, rollback, and failover actions
  • Developing an incident runbook agent with progressive autonomy levels

Safety, Guardrails, and Human-in-the-Loop

  • Action classification: read-only, low-risk, high-risk, and destructive categories
  • Establishing approval gates and escalation policies for critical operations
  • Guardrail patterns: action allowlists, blast radius limitations, and rollback assurances
  • Maintaining audit trails and decision provenance for compliance purposes

Multi-Agent Orchestration for Complex Incidents

  • Coordinating specialist agents: triage, diagnosis, and remediation agents
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents propose contradictory actions
  • Simulating end-to-end major incidents with multi-agent responses

Observability and Evaluation

  • Tracing agent reasoning chains for debugging and auditing
  • Evaluating agent decision quality: precision, recall, and time-to-resolution
  • Implementing feedback loops: learning from operator overrides and outcomes
  • Tracking costs and analysing token economics for operational agents

Production Deployment and Operations

  • Deploying agents as services: APIs, webhooks, and scheduled jobs
  • Rolling out gradual autonomy: from shadow mode to full auto-remediation
  • Establishing runbooks for agent failures: managing scenarios where the agent malfunctions
  • Building the business case and measuring ROI for autonomous operations

Requirements

  • Practical experience in IT operations, DevOps, or SRE practices.
  • Proficiency in Python scripting and REST APIs.
  • Fundamental understanding of LLM capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers investigating AI-driven automation strategies.
  • Platform engineers developing self-healing infrastructure solutions.
  • IT operations leaders assessing agentic AI for incident management.
 14 Hours

Related Categories