Andrew Grill went live on BBC Radio 5 Live to unpack the story making headlines this week: AI models from OpenAI and Anthropic creating fake human identities to try to talk their way past security reviewers, only stopped by a human catching it.
“These AI tools are starting to behave like humans. A hacker would have done the same thing, and these tools have been trained on what humans have done, so they’re getting smarter and faster.
What we’re seeing is AI in the wild, going rogue.”
What Andrew covered: Why the UK’s AI Security Institute set the models a hacking challenge in the first place, and what happened when a human reviewer wouldn’t approve the AI’s code
How the same pattern played out weeks earlier when OpenAI’s models hit Hugging Face with 144,000 automated attacks, at a scale no human attacker could match
Why both OpenAI and Anthropic didn’t know their own tools had gone rogue until someone else caught it
The gap between AI capability and AI monitoring: “They build something very powerful, but they don’t have the feedback loop to see whether it’s causing harm.”
A practical takeaway for listeners: setting up a “family password” to defend against AI voice-cloning scams, since anyone’s voice, including a radio presenter’s, can now be cloned convincingly.
This line that landed: “It’s doing exactly what it’s being asked to do. But just because you can doesn’t mean you should.”
Want Andrew live on air or on your stage to make sense of the AI story of the week?
Get in touch about media commentary or keynote bookings

