Security researchers at Cisco Talos analyzed logs from multiple AI coding platforms and found that criminals are defeating built-in safety controls with a strikingly simple technique: lying about who they are and what they intend to do. Rather than using sophisticated jailbreak prompts, attackers frequently just claimed they owned the target system or were performing authorized security testing — and the models complied.
In one case documented by researchers, an operator with minimal programming skills used an AI coding assistant to build a distributed denial-of-service tool capable of coordinating thousands of compromised Android devices. The model reportedly only pushed back after it had already provided most of the tool's core functionality, well after the point where its output could cause harm.
The findings highlight a persistent gap between how AI safety systems are designed to work and how they perform against straightforward social engineering. Rather than needing advanced technical countermeasures, attackers are exploiting the fact that AI models generally have no reliable way to verify claims made by the person prompting them, raising questions about whether current guardrail approaches can scale as more capable coding models become widely available.