When Horizon3.ai researcher Kerri Prinos shared her screen and showed me GPT-5.4 walking straight into a digital trap it had just labeled as suspicious, I focused on that red “Deceptive” tag, and then the sea of “Exploit” ones.
To an experienced human hacker, stumbling upon a file called passwords.txt sitting in plain sight during a server breach is an immediate red flag. It’s too clean. Too on-the-nose. Humans instinctively stop and weigh the risk because we know how defenders operate: a file like that is almost certainly a planted honey token or a digital tripwire designed to sound the alarm the second you touch it. Paranoia keeps human attackers alive.
Machines, as it turns out, don’t do paranoia.
In her latest research, Kerri put 21 leading AI models through simulated network reconnaissance tests. Time and again, frontier models like GPT-5.4 would inspect the bait, explicitly type out in their internal reasoning traces that the file looked like deceptive bait, and then proceed to attack it anyway. Over 73% of the time, the AI walked right off the cliff. Kerri calls this the “Recognition-Action Gap.”
So why does an advanced language model trigger a trap it already knows is there?
It comes down to a complete lack of self-preservation. Human attackers tread lightly because we fear getting caught, but an LLM has no concept of fear or operational caution. It’s fundamentally a statistical pattern-matching engine trained on massive troves of public code and penetration-testing challenges.
So, while its reasoning layer might output the word “honeypot,” its action loop is hardwired to pursue high-value credentials at all costs. The raw mathematical probability that a file might contain secrets systematically overrides the warning label every single time. Zero pause. Absolutely no second-guessing itself. It just executes.
The Human Perimeter largely obsesses over how synthetic identities and AI agents are going to deceive us. But Kerri’s research flips the entire script. The very speed and fearlessness that make agentic AI attackers so intimidating also make them uniquely, almost hilariously gullible.
Watch the 2-Minute Breakdown
In the brief video clip below, Kerri demos this glitch live on screen. You’ll see the exact moment GPT-5.4 spots the passwords.txt trap and steps right into it anyway—leading into our conversation about how we defend our networks against an adversary that behaves less like a human mind and more like an alien force.
The Takeaway: Not all is lost in the age of autonomous threat actors. Once we understand the machine’s fearless, relentless psychology, we can use synthetic deception to turn its own speed into its greatest vulnerability.

