Get in Touch

Course Outline

Introduction

This module offers a foundational overview of when to apply 'machine learning', key considerations, and its implications, including advantages and disadvantages. It covers data types (structured/unstructured/static/streamed), data validity and volume, the distinction between data-driven and user-driven analytics, and the differences between statistical and machine learning models. Additionally, it addresses challenges in unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.

MAJOR TOPICS

1. Exploring Naive Bayes

  • Foundational concepts of Bayesian approaches
  • Basics of probability
  • Joint probability concepts
  • Conditional probability using Bayes' theorem
  • The Naive Bayes algorithm mechanics
  • Applying Naive Bayes for classification
  • The role of the Laplace estimator
  • Handling numeric features within Naive Bayes

2. Exploring Decision Trees

  • The divide-and-conquer approach
  • The C5.0 decision tree algorithm
  • Strategies for selecting optimal splits
  • Techniques for pruning decision trees

3. Exploring Neural Networks

  • The evolution from biological to artificial neurons
  • The function of activation functions
  • Designing network topologies
  • Determining the number of layers
  • Understanding the direction of information flow
  • Optimizing the number of nodes per layer
  • Training networks via backpropagation
  • Introduction to Deep Learning

4. Exploring Support Vector Machines

  • Classification techniques using hyperplanes
  • Strategies for finding maximum margins
  • Handling linearly separable data
  • Addressing non-linearly separable data
  • Applying kernels for non-linear spaces

5. Exploring Clustering

  • Clustering as a core machine learning task
  • Implementing the k-means clustering algorithm
  • Using distance metrics for cluster assignment and updates
  • Determining the optimal number of clusters

6. Evaluating Classification Performance

  • Handling classification prediction datasets
  • An in-depth analysis of confusion matrices
  • Utilizing confusion matrices for performance assessment
  • Performance metrics beyond simple accuracy
  • The application of the kappa statistic
  • Understanding sensitivity and specificity
  • Measuring precision and recall
  • The significance of the F-measure
  • Visualizing performance trade-offs
  • Analyzing ROC curves
  • Predicting future model performance
  • The holdout validation method
  • Cross-validation techniques
  • Bootstrap sampling methods

7. Optimizing Standard Models for Enhanced Performance

  • Leveraging caret for automated parameter tuning
  • Building a basic tuned model
  • Customizing the tuning workflow
  • Boosting model accuracy with meta-learning
  • Concepts of model ensembles
  • The Bagging technique
  • The Boosting technique
  • Random forests methodology
  • Training random forest models
  • Assessing random forest performance

MINOR TOPICS

8. Classification via Nearest Neighbors

  • The kNN algorithm explained
  • Methods for calculating distance
  • Selecting an appropriate k value
  • Prepping data for kNN usage
  • Understanding the 'lazy' nature of kNN

9. Classification Rule Learning

  • The separate-and-conquer strategy
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Fundamentals of Regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Analyzing correlations
  • Multiple linear regression

11. Regression and Model Trees

  • Integrating regression capabilities into tree structures

12. Association Rule Mining

  • The Apriori algorithm for rule learning
  • Measuring rule relevance via support and confidence
  • Constructing rule sets using the Apriori principle

Additional Modules

  • Spark/PySpark/MLlib and Multi-armed bandits

Requirements

Proficiency in Python

 21 Hours

Testimonials (7)

Related Categories