Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon
Stanford professor and Inception co-founder Stefano Ermon explains why diffusion architecture, long successful in images and video, is poised to replace autoregressive models for text and code generation. He details how parallel token generation in diffusion models eliminates the sequential bottleneck of LLMs, dramatically improving inference speed and GPU utilization on standard hardware. The discussion covers Inception’s Mercury models for real-time voice agents, the custom software stack needed to serve diffusion-based language models at scale, and academia’s ongoing role in pursuing unconventional AI approaches. Ermon predicts that the next era of AI competition will be defined by training and inference efficiency.
Preview
4 more ideas & all timestamps
This episode is in its early-access window. The full breakdown unlocks free in about 70 hours — Pro members read everything the moment it lands.