← Home
TWIML #670 · February 5, 2024 · 55m

Reinforcement Learning in the Age of LLMs

Kamyar Azizzadenesheli from NVIDIA discusses how the reinforcement learning community is adapting to the dominance of large language models — where RL techniques enhance LLMs and where LLMs change RL.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Canon

Azizzadenesheli describes the benchmark treadmill: researchers create evaluation benchmarks that models quickly saturate. Each saturated benchmark demands a harder replacement. The field perpetually creates and destroys its own evaluation standards.

Highlights

RLHF is the most impactful application of reinforcement learning in history — yet most RL researchers did not predict it
Azizzadenesheli notes the irony: reinforcement learning from human feedback (RLHF) is the technique that made ChatGPT work, making it the most commercially impactful RL application ever. Yet most RL researchers were focused on game-playing and robotics, not language model alignment.
Was this useful?