Get in Touch
 Duration 14 hours

Course Outline

Getting Machine Learning Models Ready for Deployment

  • Packaging models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Strategies for versioning and storage

Serving Models on Kubernetes

  • Introduction to inference servers
  • Deploying TensorFlow Serving and TorchServe
  • Establishing model endpoints

Optimizing Inference Performance

  • Implementing batching strategies
  • Managing concurrent request handling
  • Tuning latency and throughput

Autoscaling ML Workloads

  • Horizontal Pod Autoscaler (HPA)
  • Vertical Pod Autoscaler (VPA)
  • Kubernetes Event-Driven Autoscaling (KEDA)

Provisioning GPUs and Managing Resources

  • Setting up GPU nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Model Rollout and Release Methodologies

  • Blue/green deployments
  • Canary rollout patterns
  • A/B testing for model assessment

Monitoring and Observability for Production ML

  • Metrics for inference workloads
  • Best practices for logging and tracing
  • Dashboards and alerting systems

Security and Reliability Factors

  • Protecting model endpoints
  • Network policies and access controls
  • Maintaining high availability

Conclusion and Future Steps

Requirements

  • Knowledge of containerized application workflows
  • Proficiency with Python-based machine learning models
  • Basic familiarity with Kubernetes concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Testimonials (3)

Related Categories