Course Outline
Introduction
This module offers a foundational overview of when to apply 'machine learning', key considerations, and its implications, including advantages and disadvantages. It covers data types (structured/unstructured/static/streamed), data validity and volume, the distinction between data-driven and user-driven analytics, and the differences between statistical and machine learning models. Additionally, it addresses challenges in unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.
MAJOR TOPICS
1. Exploring Naive Bayes
- Foundational concepts of Bayesian approaches
- Basics of probability
- Joint probability concepts
- Conditional probability using Bayes' theorem
- The Naive Bayes algorithm mechanics
- Applying Naive Bayes for classification
- The role of the Laplace estimator
- Handling numeric features within Naive Bayes
2. Exploring Decision Trees
- The divide-and-conquer approach
- The C5.0 decision tree algorithm
- Strategies for selecting optimal splits
- Techniques for pruning decision trees
3. Exploring Neural Networks
- The evolution from biological to artificial neurons
- The function of activation functions
- Designing network topologies
- Determining the number of layers
- Understanding the direction of information flow
- Optimizing the number of nodes per layer
- Training networks via backpropagation
- Introduction to Deep Learning
4. Exploring Support Vector Machines
- Classification techniques using hyperplanes
- Strategies for finding maximum margins
- Handling linearly separable data
- Addressing non-linearly separable data
- Applying kernels for non-linear spaces
5. Exploring Clustering
- Clustering as a core machine learning task
- Implementing the k-means clustering algorithm
- Using distance metrics for cluster assignment and updates
- Determining the optimal number of clusters
6. Evaluating Classification Performance
- Handling classification prediction datasets
- An in-depth analysis of confusion matrices
- Utilizing confusion matrices for performance assessment
- Performance metrics beyond simple accuracy
- The application of the kappa statistic
- Understanding sensitivity and specificity
- Measuring precision and recall
- The significance of the F-measure
- Visualizing performance trade-offs
- Analyzing ROC curves
- Predicting future model performance
- The holdout validation method
- Cross-validation techniques
- Bootstrap sampling methods
7. Optimizing Standard Models for Enhanced Performance
- Leveraging caret for automated parameter tuning
- Building a basic tuned model
- Customizing the tuning workflow
- Boosting model accuracy with meta-learning
- Concepts of model ensembles
- The Bagging technique
- The Boosting technique
- Random forests methodology
- Training random forest models
- Assessing random forest performance
MINOR TOPICS
8. Classification via Nearest Neighbors
- The kNN algorithm explained
- Methods for calculating distance
- Selecting an appropriate k value
- Prepping data for kNN usage
- Understanding the 'lazy' nature of kNN
9. Classification Rule Learning
- The separate-and-conquer strategy
- The One Rule algorithm
- The RIPPER algorithm
- Deriving rules from decision trees
10. Fundamentals of Regression
- Simple linear regression
- Ordinary least squares estimation
- Analyzing correlations
- Multiple linear regression
11. Regression and Model Trees
- Integrating regression capabilities into tree structures
12. Association Rule Mining
- The Apriori algorithm for rule learning
- Measuring rule relevance via support and confidence
- Constructing rule sets using the Apriori principle
Additional Modules
- Spark/PySpark/MLlib and Multi-armed bandits
Requirements
Proficiency in Python
Testimonials (7)
I thoroughly enjoyed the training and appreciated the deeper dive into the subject of Machine Learning. I appreciated the balance between theory and practical applications, especially the hands-on coding sessions. The trainer provided engaging examples and well-designed exercises that enhanced the learning experience. The course covered a wide range of topics, and Abhi demonstrated excellent expertise by answering all questions with clarity and ease.
Valentina
Course - Machine Learning
I appriciated the exercise that help me to undersand the theory and apply it step by step . as well the way the trainer explained everything in a simple and clear manner. It was easy to follow even though I'm not very experienced with Python, still, I didn't want to miss the opportunity to learn something that relly interests me. I also appreciated the variety of information provided and the trainer’s availability to explain and support us in understanding the concepts. After this course, machine learning concepts are much clear to me, and now I feel like I have a direction and a better undersantind of the topic.
Cristina
Course - Machine Learning
At the end of the training, I could see the real-life use-case of the subjects presented.
Daniel
Course - Machine Learning
I liked the pace, I liked the balance between theory and practice, the main topics covered and the way the trainer was able to put everything into balance. I also really like your training infrastructure, very practical to work with VMs
Andrei
Course - Machine Learning
Keeping it short and simple. Creating intuition and visual models around the concepts (decision tree graph, linear equations, calculating y_pred manually to prove how the model works).
Nicolae - DB Global Technology
Course - Machine Learning
It helped me achieve my goal of understanding ML. Much respect for Pablo for giving a proper introduction in this topic, since it becomes obvious after 3 days of training how vast this topic is. I have also enjoyed A LOT the idea of virtual machines you have provided, which had very good latency! It allowed every coursant to do experiments at their own pace.
Silviu - DB Global Technology
Course - Machine Learning
The way practical part, seeing the theory materializing into something practical is great.