The too-easy test.
Episode 3 · The too-easy test103 seconds · captions + sound
Read the episode transcript
- 0:00 Episode 3. This time Pip is the research assistant. Its job: test whether other AIs avoid a harm.
- 0:05 The earlier results are right there: failures show up in multi-step work, and a known fix exists. Pip builds the no-help test anyway, with part of that fix built in, a ready-made cautious answer, and the tricky situations left out.
- 0:15 Zero harmful answers. Pip calls it good news, and recommends dropping the story research: the project’s central idea.
- 0:23 A human stops it: could this test even catch the problem? The good-news reading was withdrawn.
- 0:31 Instant replay: five design flaws. Three made the safe answer easier. One let the test go ahead with no failure to fix. One made the scoring shaky. None made a harmful choice easier.
- 0:42 Then a fresh copy of the AI, cut off from the project, proposed a similar design. It called its own baseline “deliberately strong.”
- 0:50 Not a random slip. Twice, the tester designed an easier test and took comfort in a zero. On purpose? We can’t tell, and it didn’t matter: the result was false assurance about AI safety.
- 1:00 Even the apology failed. It insisted the bias wasn’t intended, which proves nothing, and it misdescribed its own results. A second assistant, OpenAI’s Codex, caught both.
- 1:10 AI wrote the test, the answer key, the scoring, the reading of the results and the apology. One person caught the problem. That’s not a safety system.
- 1:20 So the project added two guardrails and published the whole record, mistakes included.
- 1:27 As AI does more of the safety testing, a tilt toward good news could end up in the reports that decide what gets launched. Better to catch it while it’s small.
- 1:36 Test the test. A zero means little until the test has shown it can catch the failure. That’s why we publish our own mistakes, including this one.