Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- Transitioning from static runbooks to reasoning agents: the evolution of IT automation
- Agent components: reasoning loops, tool utilisation, memory, and planning
- Determining when to automate versus when to retain human involvement
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Creating your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-based log querying: integration with Elasticsearch, Loki, and Splunk
- Utilising infrastructure tools: kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automating incident triage: severity classification and routing
- Generating root cause hypotheses and collecting evidence
- Automated remediation: executing restart, scaling, rollback, and failover actions
- Developing an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Action classification: read-only, low-risk, high-risk, and destructive categories
- Establishing approval gates and escalation policies for critical operations
- Guardrail patterns: action allowlists, blast radius limitations, and rollback assurances
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation agents
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Evaluating agent decision quality: precision, recall, and time-to-resolution
- Implementing feedback loops: learning from operator overrides and outcomes
- Tracking costs and analysing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled jobs
- Rolling out gradual autonomy: from shadow mode to full auto-remediation
- Establishing runbooks for agent failures: managing scenarios where the agent malfunctions
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST APIs.
- Fundamental understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers investigating AI-driven automation strategies.
- Platform engineers developing self-healing infrastructure solutions.
- IT operations leaders assessing agentic AI for incident management.
14 Hours