← Home
No Priors · June 26, 2026 · 36m
Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
OpenAI research scientist Noam Brown joins Sarah Guo to discuss his essay on why AI benchmarks are broken when models scale test-time compute. They explore how today's models can reason for weeks or months on complex tasks like math conjectures and poker solving, and why static benchmark grids fail to capture this capability. The conversation covers safety implications of variable compute budgets, the limits of recursive self-improvement, and the future of multi-agent collaboration.
This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.
Canon
•
Over-optimizing models to score well on specific benchmarks leads to performance that does not generalize to real-world usefulness, mirroring Goodhart's Law.
Highlights
•
•
Long-Horizon Reasoning Unlocks Impossible Tasks08:34
With proper scaffolding and test-time compute budgets stretching to days or months, models can reason through problems that previously required human-level insight, such as disproving math conjectures and building poker solver bots.Was this useful?