← Home
No Priors · June 26, 2026 · 36m

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

OpenAI research scientist Noam Brown joins Sarah Guo to discuss his essay on why AI benchmarks are broken when models scale test-time compute. They explore how today's models can reason for weeks or months on complex tasks like math conjectures and poker solving, and why static benchmark grids fail to capture this capability. The conversation covers safety implications of variable compute budgets, the limits of recursive self-improvement, and the future of multi-agent collaboration.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Canon

Over-optimizing models to score well on specific benchmarks leads to performance that does not generalize to real-world usefulness, mirroring Goodhart's Law.

Highlights

Static Benchmarks Don't Capture Variable Compute Capability01:23
Traditional AI benchmarks assume a fixed inference budget, but models that can dynamically allocate more time to 'think' achieve far better performance on complex tasks.
Long-Horizon Reasoning Unlocks Impossible Tasks08:34
With proper scaffolding and test-time compute budgets stretching to days or months, models can reason through problems that previously required human-level insight, such as disproving math conjectures and building poker solver bots.
Safety Evaluations Must Scale with Compute Budget11:26
Existing safety frameworks assess models at a fixed compute level, but dangerous capabilities can emerge when models are given more time to think, requiring evals that track how risk scales with inference budget.
Was this useful?