← Home
TWIML · August 12, 2026 · 56m

Why Image Generation Needs More Than Bigger Models with Fatih Porikli

Text-to-image models have become remarkably good at producing realistic images, but Fatih Porikli argues that realism is not the same as correctness — models still fail at several distinct people, precise compositions, and local high-resolution generation. The Qualcomm Vice President of Technology walks through approaches his team presented at CVPR: better training objectives for controllability, separating scene planning from rendering, and generating 16-megapixel images efficiently on edge devices. The conversation also covers eliminating visible artifacts in AI-powered editing, reinforcement learning for image generation, and agentic pipelines. Porikli maps what remains unsolved in generative vision and where the next phase of progress is likely to come from.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Novel

Better Training Objectives Improve Controllability
Porikli discusses how better training objectives can improve controllability in image generation.
Separate Scene Planning From Rendering
Porikli presents a two-stage approach that separates scene planning from rendering for more reliable image generation.
16-Megapixel Generation on Edge Devices
Porikli covers techniques for generating 16-megapixel images efficiently on edge devices.
Eliminating Artifacts in AI Image Editing
Porikli covers new methods for eliminating the visible artifacts that often appear in AI-powered image editing.
RL and Agentic Pipelines as the Next Phase
The conversation turns to reinforcement learning for image generation and agentic pipelines as directions for the next phase of generative vision.

Highlights

Realism Is Not Correctness
Porikli frames the episode around a core distinction: realism is not the same as correctness in text-to-image generation.
Was this useful?