17 juli 2026
79 min
David Manheim is head of methodology at AI Evaluation Consensus. He joins the podcast to discuss how AI evaluations can become more reliable, transparent, and useful for decisions. We cover common failures such as unclear reporting, training to the test, benchmark saturation, and models changing behavior when they know they are being tested. The conversation also examines real-world tests, biosecurity, persuasion, forecasting, human oversight, and why even “normal” AI progress could be disruptive.
LINKS:
CHAPTERS:
(00:00) Episode Preview
(01:04) Evaluation consensus project
(07:01) Evaluation awareness challenges
(12:28) Reporting capabilities clearly
(19:38) Benchmarks beyond humans
(29:52) Proxies and biosecurity
(42:01) Persuasion and democracy
(53:59) Forecasting with AI
(01:08:44) Oversight and disruption
(01:16:42) Supporting better evals
PRODUCED BY:
SOCIAL LINKS:
Website: https://podcast.futureoflife.org
Twitter (FLI): https://x.com/FLI_org
Twitter (Gus): https://x.com/gusdocker
LinkedIn: https://www.linkedin.com/company/future-of-life-institute/
YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/
Apple: https://geo.itunes.apple.com/us/podcast/id1170991978
Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP
Lyssna på fler avsnitt från
Future of Life Institute Podcast
Visar 1–10 av 271 avsnitt
30 juni 2026
44 min
25 juni 2026
64 min
12 juni 2026
68 min
26 maj 2026
74 min
11 maj 2026
96 min
7 maj 2026
67 min
29 april 2026
84 min
17 april 2026
54 min
2 april 2026
56 min
20 mars 2026
72 min