← Home
TWIML · October 14, 2024 · 52m

How Do We Actually Evaluate AI Safety?

A critical examination of AI safety evaluation methodologies: red-teaming, benchmarks, and stress testing — what they actually measure, what they miss, and why the field lacks agreed-upon standards.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Canon

Charrington identifies the structural problem: most AI safety research is funded by the same companies whose products it evaluates. This funding environment shapes which safety questions get asked (technically tractable ones) and which get ignored (ones that might slow product development).

Highlights

Current AI safety evaluations test for known risks but cannot detect unknown risks — we are looking under the lamppost because that is where the light is
The discussion reveals that AI safety evaluations test for risks that evaluators can imagine and formalize into test cases. Novel risks — failure modes that no one has anticipated — are by definition untestable. Current evaluations provide confidence about known risks while leaving unknown risks completely unaddressed.
Was this useful?