← Home
No Priors · May 1, 2026 · 42m
Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
Tuhin Srivastava, CEO of Baseten, joins Sarah Guo and Elad Gil to discuss the explosive growth in AI inference demand and Baseten's role as the inference cloud. He explains why the application layer will persist as companies encode unique user data into custom models and post-train on proprietary signals. The conversation covers GPU supply constraints, Baseten's approach to multi-cloud orchestration across 18 clouds and 90 clusters, the shift toward longer GPU contracts, and the operational challenges of scaling an inference platform. Srivastava also shares his views on Chinese AI models, the dominance of custom inference workloads, and the future of multichip and concierge-like AI services.
This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.
Canon
•
Making inference cheaper and faster leads to dramatically higher consumption, not less, mirroring a well-known economic pattern.
Highlights
•
The Application Layer's Moat: Proprietary Data and Custom Models00:01:55
Tuhin Srivastava argues that companies with unique user signals can build durable value by encoding that data into workflows and post-training specialized models, giving them an edge over generic AI solutions.•
Editorial
Misc
✧Srivastava describes a 'Concierge Everything' future where AI services become proactive assistants.
Was this useful?