Text-to-image models have become remarkably good at producing realistic images, but Fatih Porikli argues that realism is not the same as correctness — models still fail at several distinct people, precise compositions, and local high-resolution generation. The Qualcomm Vice President of Technology walks through approaches his team presented at CVPR: better training objectives for controllability, separating scene planning from rendering, and generating 16-megapixel images efficiently on edge devices. The conversation also covers eliminating visible artifacts in AI-powered editing, reinforcement learning for image generation, and agentic pipelines. Porikli maps what remains unsolved in generative vision and where the next phase of progress is likely to come from.
This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.
Novel
•
Better Training Objectives Improve Controllability
Porikli discusses how better training objectives can improve controllability in image generation.