Get in Touch

Course Outline

Introduction to Application Performance Monitoring

  • Grasping the concept of Application Performance Management (APM) and its significance in modern operations
  • The interplay between application performance, availability, reliability, and customer satisfaction
  • Essential performance indicators and service-level objectives
  • Recognizing common causes of application performance decline
  • The application monitoring lifecycle: observation, analysis, diagnosis, remediation, and optimization
  • The role of New Relic in achieving full-stack observability

New Relic Features and Architecture

  • An overview of the New Relic platform and its primary functions
  • Understanding the architectural design and data flow within New Relic
  • Components such as agents, collectors, and telemetry data streams
  • An introduction to metrics, events, logs, traces, and error tracking
  • Coverage of APM, browser monitoring, infrastructure oversight, and database monitoring
  • Understanding key entities, services, applications, and workloads
  • Foundations of distributed tracing and service interdependencies
  • Concepts related to data retention, querying, and visualization

Navigating the New Relic Interface

  • Exploring the New Relic platform and primary dashboards
  • Managing applications, services, hosts, and various entities
  • Reviewing performance summaries and overall application health
  • Utilizing charts, tables, filters, and time range selections
  • Searching and interpreting telemetry data
  • Customizing dashboards and user views
  • Creating effective operational and performance dashboards
  • Using New Relic to transition from high-level symptoms to detailed diagnostics

Setting Up and Configuring New Relic Agents

  • Understanding the agent architecture and supported environments
  • Installing New Relic agents on application servers
  • Configuring agents for specific application monitoring needs
  • Strategies for instrumentation, including automatic and manual methods
  • Setting up browser and end-user monitoring capabilities
  • Verifying agent installation and telemetry data collection
  • Managing configuration and environment-specific settings
  • Troubleshooting agent deployment and data collection challenges
  • Best practices for secure and sustainable agent deployment

Measuring Application Performance from the End-User View

  • The principles of Real User Monitoring (RUM)
  • Tracking page-load times and application response speeds
  • Monitoring browser performance and user interaction patterns
  • Identifying slow pages, transactions, and user journeys
  • Analyzing performance variations by geography and device type
  • Linking end-user experience with backend application performance
  • Spotting performance issues that directly impact customer satisfaction
  • Using performance data to prioritize optimization efforts

Reading and Understanding Instrumentation Data

  • Understanding transaction traces and application workflow
  • Interpreting data on response time, throughput, and error rates
  • Analyzing transaction breakdowns and specific performance segments
  • Evaluating external services and their dependencies
  • Reviewing application errors and associated error traces
  • Locating bottlenecks within instrumentation data
  • Using traces to track requests across various application components
  • Combining metrics, events, logs, and traces for root-cause analysis
  • Practical exercises in interpreting application telemetry

Measuring Application Resources and Infrastructure

  • Monitoring resource utilization across applications and servers
  • Understanding CPU, memory, disk, and network performance metrics
  • Identifying resource saturation and capacity constraints
  • Correlating infrastructure metrics with application response times
  • Monitoring application processes and active workloads
  • Spotting resource-intensive transactions
  • Investigating performance drops due to infrastructure limitations
  • Setting performance baselines and detecting anomalies

Monitoring and Notifications

  • Core concepts behind New Relic alerting
  • Defining alert conditions and setting thresholds
  • Creating alerts for performance and availability metrics
  • Tracking error rates, response times, throughput, and resource usage
  • Designing effective and actionable alert policies
  • Configuring notification channels and incident workflows
  • Minimizing alert noise to avoid unnecessary notifications
  • Understanding incident management and issue correlation
  • Testing and verifying alert configurations
  • Best practices for proactive application monitoring

Monitoring Database Operations

  • The link between database performance and overall application performance
  • Tracking database calls and query activity
  • Identifying slow-running database operations
  • Analyzing database response times
  • Detecting inefficient or resource-heavy queries
  • Correlating database activities with application transactions
  • Investigating application bottlenecks related to databases
  • Using performance insights to improve query response times
  • Practical exercises in diagnosing database performance issues

Reporting and Visualizing Application Performance

  • Creating insightful performance reports
  • Building dashboards for development, operations, and management teams
  • Selecting the right metrics for different stakeholders
  • Visualizing availability, response time, throughput, and errors
  • Tracking performance trends over extended periods
  • Comparing performance across different environments
  • Presentation technical metrics as business-relevant insights
  • Establishing baselines and reporting against specific objectives

Analyzing and Optimizing Application Performance

  • Establishing a structured approach to performance analysis
  • Identifying bottlenecks and abnormal system behavior
  • Analyzing transaction response times and throughput
  • Comparing current performance against historical baselines
  • Correlating multiple data sources during investigations
  • Prioritizing issues based on user and business impact
  • Identifying opportunities for application optimization
  • Validating improvements using New Relic data
  • Hands-on performance analysis exercises

Troubleshooting API and Service Issues

  • Monitoring APIs and external service dependencies
  • Measuring API response times, throughput, and error rates
  • Identifying slow or unreliable API endpoints
  • Diagnosing timeout and connectivity problems
  • Analyzing failed API transactions
  • Using traces to spot bottlenecks in distributed services
  • Correlating API issues with downstream dependencies
  • Determining the root cause of API performance drops
  • Developing and validating remediation strategies

Distributed Tracing and End-to-End Troubleshooting

  • Understanding distributed applications and service dependencies
  • Foundations of distributed tracing concepts
  • Tracking requests across multiple application services
  • Identifying latency introduced by individual services
  • Analyzing communication between services
  • Detecting failures across distributed components
  • Correlating traces with logs, errors, and infrastructure metrics
  • Conducting end-to-end root-cause analysis
  • Practical troubleshooting scenarios in a live-lab environment

Querying and Analyzing New Relic Data

  • Introduction to querying telemetry data in New Relic
  • Understanding the New Relic Query Language (NRQL)
  • Writing queries to investigate performance issues
  • Filtering and aggregating metrics and events
  • Analyzing response times, errors, throughput, and transaction data
  • Creating custom visualizations from query outputs
  • Using queries to support troubleshooting and reporting needs
  • Building reusable queries and dashboards
  • Practical NRQL exercises

Integrating New Relic with Third-Party Tools and Services

  • Overview of available New Relic integrations
  • Connecting New Relic with infrastructure and cloud platforms
  • Linking monitoring data with collaboration and incident management tools
  • Understanding integration workflows and data exchange
  • Configuring notifications and external service links
  • Leveraging integrations for DevOps and incident response
  • Best practices for maintaining reliable monitoring integrations

Practical Troubleshooting Workshop

  • Investigating a simulated application performance incident
  • Identifying symptoms from end-user performance data
  • Analyzing application transactions and error logs
  • Investigating infrastructure and database performance
  • Tracing API and external service dependencies
  • Correlating metrics, events, logs, and traces
  • Identifying the most probable root cause
  • Developing and validating a remediation approach
  • Configuring alerts to prevent recurrence
  • Documenting findings and communicating business impact

Monitoring Best Practices and Operational Recommendations

  • Designing an effective New Relic monitoring strategy
  • Selecting meaningful performance and availability metrics
  • Setting baselines and service-level objectives
  • Reducing excessive monitoring noise
  • Developing effective alerting and escalation practices
  • Maintaining consistent monitoring across development, testing, and production
  • Using observability data to drive continuous improvement
  • Translating technical data into actionable business insights

Summary and Conclusion

  • Review of New Relic architecture and core capabilities
  • Recap of application, infrastructure, database, API, and end-user monitoring
  • Summary of troubleshooting and root-cause analysis techniques
  • Review of alerting, dashboards, reporting, and integrations
  • Final hands-on performance investigation
  • Discussion of real-world implementation scenarios
  • Questions and answers session
  • Recommended next steps for applying New Relic in production

Requirements

  • A foundational grasp of application infrastructure principles
  • Familiarity with the Linux command line interface

Target Audience

  • Software Developers
  • DevOps Engineers
  • Quality Assurance/Test Engineers
  • System Administrators
  • Solution Architects
 28 Hours

Related Categories