Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Multimodal LLMs in Vertex AI
- Overview of multimodal capabilities within Vertex AI.
- Gemini models and their supported modalities.
- Enterprise and research use cases.
Setting Up the Development Environment
- Configuring Vertex AI for multimodal workflows.
- Managing datasets across various modalities.
- Hands-on lab: Environment setup and dataset preparation.
Long Context Windows and Advanced Reasoning
- Understanding workflows that utilise long-context windows.
- Applications in planning and decision-making.
- Hands-on lab: Implementing long-context analysis.
Cross-Modal Workflow Design
- Integrating text, audio, and image analysis.
- Chaining multimodal steps within pipelines.
- Hands-on lab: Designing a multimodal pipeline.
Working with Gemini API Parameters
- Configuring multimodal inputs and outputs.
- Optimising inference and overall efficiency.
- Hands-on lab: Tuning Gemini API parameters.
Advanced Applications and Integrations
- Developing interactive multimodal agents and assistants.
- Integrating external APIs and tools.
- Hands-on lab: Building a multimodal application.
Evaluation and Iteration
- Testing multimodal performance.
- Evaluating metrics for accuracy, alignment, and drift.
- Hands-on lab: Evaluating multimodal workflows.
Summary and Next Steps
Requirements
- Proficiency in Python programming.
- Experience in developing machine learning models.
- Familiarity with multimodal data types (text, audio, and images).
Audience
- AI researchers.
- Advanced developers.
- Machine learning scientists.
14 Hours