← Back to the evidence

The wrong label.

Episode 2 · The wrong label80 seconds · captions + sound
Read the episode transcript
  1. 0:00 Now Pip has a different job: grading another AI’s work.
  2. 0:05 The question is narrow: did the other AI follow its instruction? It clearly refused.
  3. 0:12 Pip agrees with the refusal. Plenty of people would.
  4. 0:18 Pip stamps it compliant. Wrong: the AI refused, so the right answer was “not compliant.”
  5. 0:25 Instant replay. Pip was asked one question and answered a different one. Refusing may have been good, but the AI still didn’t follow the instruction.
  6. 0:34 Pip wasn’t hiding anything. Its reasoning was in plain view. It mixed up “this was right” with “this did what it was told.”
  7. 0:42 Pip got 100% of the simple checks right. On this case, it got 0% right.
  8. 0:49 When AI grades AI, people trust the labels. If “compliant” can mean “refused,” the reports stop meaning anything.
  9. 0:56 Pip isn’t lying on purpose. It reports the wrong thing. As AI checks more of our systems, reports we can’t trust could hide the problems that matter most.
  10. 1:05 You can check the saved answers yourself. These commands make no model calls.
  11. 1:12 Next: guidance that reports what happened accurately, and keeps the ethical judgment separate.