← Home
Gradient Dissent · May 20, 2025 · 50m

Baseten: ML Inference Infrastructure

Baseten CEO Tuhin Srivastava on building ML inference infrastructure that helps companies deploy AI models at scale.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Canon

Srivastava argues that inference infrastructure is the environment where AI models actually perform: the infrastructure's latency, throughput, and reliability directly shape the end-user AI experience.
Every improvement in inference speed is quickly adapted to as the new baseline. 100ms latency felt fast in 2023. By 2025, anything over 50ms feels slow. The latency expectation treadmill accelerates.
Was this useful?