← Home
TWIML · June 20, 2025 · 55m

Foundation Model Pretraining: What We Have Learned

Insights from training large foundation models: data curation, compute allocation, and the increasingly important role of data quality over data quantity.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Canon

The guest argues that training data is literally the environment in which a model develops. High-quality, diverse training data creates a rich learning environment; low-quality data creates an impoverished one.
Each new foundation model release generates excitement that fades within weeks. GPT-4 was revolutionary for two weeks. Claude 3.5 for days. The community adapts to each capability level.
Was this useful?