Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Solutions
- Key concepts and advantages of AIOps.
- The role of Prometheus and Grafana within the observability ecosystem.
- The position of machine learning in AIOps: comparing predictive and reactive analytics.
Configuring Prometheus and Grafana
- Installation and setup of Prometheus for time series data collection.
- Designing Grafana dashboards driven by real-time metrics.
- Investigating exporters, relabeling processes, and service discovery.
Preparing Data for Machine Learning
- Extraction and transformation of Prometheus metrics.
- Structuring datasets suitable for anomaly detection and forecasting tasks.
- Utilizing Grafana’s transformation features or Python-based data pipelines.
Utilizing Machine Learning for Anomaly Detection
- Core machine learning models for outlier identification (e.g., Isolation Forest, One-Class SVM).
- Model training and evaluation using time series data.
- Visualizing detected anomalies within Grafana dashboards.
Metric Forecasting via Machine Learning
- Development of basic forecasting models (including ARIMA, Prophet, and LSTM introductions).
- Anticipating system load or resource consumption patterns.
- Leveraging predictions for proactive alerting and scaling decisions.
Machine Learning Integration with Alerting and Automation
- Creating alert rules driven by machine learning outputs or specific thresholds.
- Implementing Alertmanager and managing notification routing.
- Initiating scripts or automation workflows upon anomaly detection.
Scaling and Operationalizing AIOps
- Incorporating external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace).
- Managing machine learning models within observability pipelines.
- Best practices for implementing AIOps at scale.
Conclusion and Recommendations for Next Steps
Requirements
- A solid grasp of system monitoring and observability principles.
- Practical experience with either Grafana or Prometheus.
- Knowledge of Python and fundamental machine learning concepts.
Target Audience
- Observability engineers.
- Members of infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).