← Home
No Priors · September 18, 2026 · 38m

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

Stanford professor and Inception co-founder Stefano Ermon explains why diffusion architecture, long successful in images and video, is poised to replace autoregressive models for text and code generation. He details how parallel token generation in diffusion models eliminates the sequential bottleneck of LLMs, dramatically improving inference speed and GPU utilization on standard hardware. The discussion covers Inception’s Mercury models for real-time voice agents, the custom software stack needed to serve diffusion-based language models at scale, and academia’s ongoing role in pursuing unconventional AI approaches. Ermon predicts that the next era of AI competition will be defined by training and inference efficiency.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Preview

•
Why Diffusion Beats Autoregressive for AI Inference05:59
Ermon explains that autoregressive LLMs generate text one token at a time, creating a sequential bottleneck that underutilizes GPU parallelism, whereas diffusion models generate all tokens in parallel, offering superior inference scaling and hardware efficiency.
•
Inception's Mercury Models and Voice Agents13:19
Inception has developed Mercury, a family of diffusion-based language models optimized for real-time voice agent applications, and built a custom inference stack to serve them at scale.

4 more ideas & all timestamps

This episode is in its early-access window. The full breakdown unlocks free in about 70 hours — Pro members read everything the moment it lands.

Read it now with Pro$10/mo · founding $96/yr
Was this useful?