Engineering GPU performance: Design choices for training & inference

  • Watch on-demand now
  • The Compute, Storage & Networking Channel
Speakers

This episode is available to stream on-demand.

This episode discusses the technical nuances of GPU performance and system design for AI and HPC. Expert speakers will compare hosted cloud and on-prem strategies for both training and large-scale inference, examining elasticity, queue times, latency, and cost predictability, while exploring how software-stack maturity and orchestration shape throughput and iteration speed. Join this episode to uncover:

  • What factors are driving hybrid inference strategies?
  • Measuring and monitoring GPU performance
  • Thermal and serviceability trade-offs for outcome-driven GPU systems
  • GPU optimization: monitoring, placement, and ROI

Related Episodes