Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction and Diagnostic Foundations
- An overview of common failure modes in LLM systems and specific issues related to Ollama.
- Establishing reproducible experiments and controlled testing environments.
- Essential debugging tools: local logs, request/response captures, and sandboxing.
Reproducing and Isolating Failures
- Methods for creating minimal failing examples and initial test seeds.
- Distinguishing between stateful and stateless interactions to isolate context-dependent bugs.
- Managing determinism, randomness, and controlling non-deterministic behaviors.
Behavioral Evaluation and Metrics
- Quantitative measures: accuracy, ROUGE/BLEU variations, calibration, and perplexity indicators.
- Qualitative assessments: human-in-the-loop scoring and the design of evaluation rubrics.
- Task-specific fidelity verification and definition of acceptance criteria.
Automated Testing and Regression
- Unit tests for prompts and components, as well as scenario and end-to-end tests.
- Building regression suites and establishing baselines with golden examples.
- Integrating CI/CD for Ollama model updates and implementing automated validation gates.
Observability and Monitoring
- Structured logging, distributed tracing, and the use of correlation IDs.
- Key operational metrics: latency, token usage, error rates, and quality signals.
- Configuration of alerting, dashboards, and SLIs/SLOs for model-driven services.
Advanced Root Cause Analysis
- Tracing through graphed prompts, tool invocations, and multi-turn conversation flows.
- Conducting comparative A/B diagnostics and ablation studies.
- Data provenance tracking, dataset debugging, and resolving dataset-induced failures.
Safety, Robustness, and Remediation Strategies
- Mitigation techniques: filtering, grounding, retrieval augmentation, and prompt scaffolding.
- Implementing rollback, canary, and phased rollout patterns for model updates.
- Conducting post-mortems, analyzing lessons learned, and establishing continuous improvement cycles.
Summary and Next Steps
Requirements
- Substantial experience in developing and deploying LLM applications.
- Proficiency with Ollama workflows and model hosting practices.
- Proficiency in Python, Docker, and fundamental observability tools.
Target Audience
- AI Engineers
- ML Ops Professionals
- QA teams managing production LLM systems