This intensive three-day workshop is dedicated to constructing and refining high-performance data processing workloads within Kubernetes environments, leveraging PySpark, Pandas and Polars.
Attendees will gain a working knowledge of how Spark applications operate on Kubernetes, along with the impact that application-level configuration choices have on overall performance, scalability, resource utilisation and operational costs. The curriculum explores critical optimisation domains such as executor sizing, memory distribution, dynamic allocation, partitioning methods, shuffle operations, mitigating small-file issues and implementing efficient Parquet handling.
Additionally, the session tackles frequent difficulties associated with Pandas, such as memory constraints and out-of-memory errors, while presenting Polars as a high-performance alternative for specific processing tasks. Through practical exercises, participants will learn to diagnose performance and memory challenges, evaluate various configuration strategies and implement optimisation techniques for realistic ETL and machine learning contexts.
Throughout the workshop, the focus remains on practical decision-making: identifying bottlenecks, choosing the right tools, configuring Spark effectively and balancing performance against infrastructure resource usage and expenditure.
Read more...