← Home
The Vergecast · June 25, 2026 · 26m

How to train your data

The Vergecast discusses training data—the raw material of AI. Alex Reisner, a staff writer at The Atlantic, explains how AI companies gather massive datasets from books, blogs, YouTube, and Reddit, why they prefer to keep that data secret, and whether the practice can ever be fair. Reisner also highlights specific findings, including that over 15 million YouTube videos have been scraped and millions of songs incorporated into AI-generated music.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Highlights

The composition of AI training data
Training data for large AI models includes books, blog posts, YouTube videos, and Reddit comments, among other sources.
At least 15 million YouTube videos used for AI training
Researchers found that at least 15 million YouTube videos have been scraped and used as training data for AI models.
Common Crawl's controversial role
Common Crawl, a nonprofit that archives the web, is a primary source of training data for AI companies, but its practices raise significant ethical questions.
Millions of songs in AI-generated music
AI music generators have been trained on millions of songs, often without artist consent or compensation.
The fairness of the training data exchange
Reisner asks whether the current system, where AI companies take vast amounts of publicly available data for free, could ever be considered a fair trade for the individuals and industries whose work is used.
Was this useful?