This episode is available to stream on-demand.
This episode discusses the technical nuances of GPU performance and system design for AI and HPC. Expert speakers will compare hosted cloud and on-prem strategies for both training and large-scale inference, examining elasticity, queue times, latency, and cost predictability, while exploring how software-stack maturity and orchestration shape throughput and iteration speed. Join this episode to uncover:
- What factors are driving hybrid inference strategies?
- Measuring and monitoring GPU performance
- Thermal and serviceability trade-offs for outcome-driven GPU systems
- GPU optimization: monitoring, placement, and ROI