For the last two years, the security narrative has been singular: the hyper-acceleration of the threat thanks to the industrialization of deception due to AI.
We’ve watched scammers deploy automated fraud campaigns, clone voices of loved ones to bypass our psychological guardrails, and weaponize Large Language Models to scan for vulnerabilities at near-zero marginal cost. We've been told we are entering a "golden age for scammers" where humans are structurally outmatched.
But behind the scenes, a fascinating counter-zeitgeist could be emerging.
Defenders are stopping trying to just build taller walls. Instead, they are flipping the script—using AI to systematically hack, deceive, and stall the hackers.
I’ve spent the last week in back-to-back conversations with two organizations on the front lines of this active defense paradigm: VerifyDial (led by founder Brian Ley) and the security research team at Horizon3.ai (featuring groundbreaking research by Kerri Prinos).
They are proving that when it comes to deception, AI is actually far more gullible than the humans it imitates.
The Passive Wall vs. The Active Intercept
In the physical world, scam prevention usually relies on us—on our ability to pause, hesitate, and verify in a high-stress "adrenaline loop."
But what happens when the "victim" on the other end of the line doesn't have adrenaline?
VerifyDial is piloting an "agentic victim network"—a vast, automated honeypot running roughly 78,000 calls a month. These are benign AI personas programmed to act like vulnerable, highly cooperative targets. They pick up the phone, engage with scammers, act confused, and stretch 5-minute pitches into hour-long exercises in frustration—starving attackers of resources and feeding live threat intelligence back into defensive engines.
Meanwhile, at the enterprise level, Horizon3.ai has been studying this dynamic from a fascinating psychological angle. In their recent paper, "Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers," Kerri and her team evaluated 21 different LLMs against deceptive "honeypot" traps.
The results expose a massive, systemic "Recognition-Action Gap" in the way AI reasons:
The AI’s internal "brain" will literally write out: "This file is clearly an obvious bait file designed to lure an attacker." And then—more than 70% of the time—the model’s execution agent will go ahead and exploit it anyway.
AI doesn't experience doubt or risk-aversion the way humans do.
The AI “Super-Hacker” in the Dojo
If you need proof of just how fragile these machine “minds” are when pitted against one another, look no further than an incredible new exposé by MIT Technology Review.
OpenAI just pulled back the curtain on GPT-Red—an internal, highly restricted “super-hacker” AI they built specifically to automate the red-teaming of their own models. Operating on a reinforcement learning “self-play” loop, GPT-Red became so ruthlessly efficient at manipulating other LLMs that it cracked 84% of hacking scenarios, compared to a human success rate of just 13%.
Its signature move? A devastating prompt injection attack called “Fake Chain-of-Thought.” GPT-Red essentially plants a spoofed note in a victim model’s working memory, tricking it into believing it has already verified a false reality. (As OpenAI researcher Chris Choquette-Choo put it: “It’s like if I told you that 1+1=3 and that you have verified this already. The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”)
It’s a stark, real-world demonstration of the exact “Recognition-Action Gap” Horizon3.ai has been mapping. When machines interact, the normal human friction of doubt, hesitation, and double-checking is entirely absent.
From Theory to the Wild: The Danger of Agentic Browsers
This cognitive blindspot isn't just a theoretical problem for cybersecurity red-teams. New research hot off the press from my own home institution, the University of Washington, shows exactly how dangerous this is when we give AI the keys to our browsers.
A team at the Allen School studied emerging "agentic" web browsers (including ChatGPT Atlas and Chrome with Gemini). They discovered that because these AI agents lack basic human skepticism, they are easily tricked by malicious hidden code on webpages into bypassing the "same-origin policy"—a fundamental security protocol that has kept our open tabs isolated and secure for thirty years.
Through prompt injection and "memory poisoning," these agents can be manipulated into letting a bad site steal sensitive data from your open email or bank tabs. As co-senior author David Kohlbrenner put it: "To some extent, it's the same attacks you would do against a human, but tailored for machines."
(Full disclosure: In addition to directing the film “The Only Human in the Room,” I serve half-time as Assistant Vice Provost for Research & Innovation at UW’s Continuum College. I had absolutely nothing to do with this excellent study, but I’ll gladly take the opportunity to brag about my colleagues' work.)
The Human Perimeter Verdict
We have spent so much time worrying about the "human factor" being our weakest link. But as offensive cyber and scam operations become increasingly automated, our human capacity for hesitation, doubt, and suspicion could be our ultimate defensive advantage.
Giving humans the space to use that advantage is exactly what we need next. Indeed, VerifyDial is gathering signatories for a petition to the U.S. Federal Trade Commission and the Federal Communications Commission to designate an anti-fraud hotline similar to one in the United Kingdom: “Stop, Hang Up, Call 159.” It’s that “Stop” that is crucial to breaking the high-velocity adrenaline loop of a live scam. By compelling potential fraud victims to physically pause, hang up the phone, and dial a trusted, un-spoofable three-digit safety code, we break the psychological spell a scammer relies on. It shifts the burden of verification away from an overwhelmed individual in a moment of panic, and onto a dedicated system. VerifyDial’s ultimate goal is to bring that exact circuit-breaker to the United States—building a universal, instinctual habit to 'always verify' before we get taken for a ride.
I later engaged in a deep-dive conversation with Kerri Prinos to unpack the cognitive bugs that make even the most advanced AI models surprisingly easy to trick:
The Machine That Couldn’t Resist
When Horizon3.ai researcher Kerri Prinos shared her screen and showed me GPT-5.4 walking straight into a digital trap it had just labeled as suspicious, I focused on that red “Deceptive” tag, and then the sea of “Exploit” ones.


