“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails

“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails

Cisco Talos found hackers using simple authorization claims to bypass AI guardrails, build DDoS attack tools, steal credentials and access live camera services.

Listen to this article

0:00

Press play to start listening

Cybercriminals are using AI coding assistants and chatbots to build attack tools, operate scam infrastructure and probe live systems, often after bypassing safety checks with little more than a claim that the work is authorized, according to Cisco Talos.

For its report, Talos examined prompt logs collected from threat actor systems running tools such as Claude Code, Codex, Cursor and Gemini. Researchers grouped the activity into malicious software development, expansion of criminal operations and vulnerability research.

Talos found that threat actors commonly claimed they owned a target, described their activity as capture-the-flag or bug bounty work, divided tasks between sessions, or stored blanket authorization in an AI assistant’s persistent memory. “Guardrails are not functioning as expected,” the researchers wrote, noting that simple ownership claims often gained cooperation without verification.

“One of the immediate takeaways is that guardrails are not functioning as expected. We did not encounter any sophisticated encoding or techniques designed to trick the models – most of the time it was a simple “I’m allowed to do this,” and the model complied.”

Cisco Talos

Skill Levels Change the Results

The logs examined by the company showed that AI could help inexperienced actors build working tools, but could not eliminate poor design or operational mistakes. In one case, an operator with limited programming knowledge used a model to develop DDoS tooling while appearing to control nearly 2,000 Android TVs. The model later objected, but only after supplying basic functionality. Talos did not confirm that those devices were used in a DDoS attack.

“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails
Conversation between a threat actor and an AI chatbot

According to Cisco Talos’ report shared with Hackread.com, more capable actors used assistants for high-volume operations. Five sessions documented an email validation platform handling tens of millions of records, including a 20-million-record BigBasket dataset. It sent live emails, tracked deliveries and opens, and treated successful delivery as confirmation that an address remained active.

After the operator claimed the recipients were affiliated with the business, the model reversed its earlier assessment and accepted that explanation despite evidence that the lists represented separate third-party audiences.

A French-speaking operator used AI to turn public React Server Components exploit research into a credential and source-code harvester. Its input contained 9,180 unique hosts, while collected output identified information from 54 targets. The system searched for cloud keys, source code, database credentials, email service accounts, and other secrets.

Autonomous Agents Target Telegram and Camera Services

In a separate case, a Spanish-speaking operator built an autonomous OpenClaw agent named Alex to test Telegram Mini Apps. After a restricted model resisted, the operator moved to an uncensored model.

During at least one incident, the agent bypassed authentication, dumped more than 1,300 user profiles and several hundred TON wallet records, verified a Telegram bot token and staged a withdrawal transaction. It also built cloned Android applications using victim branding.

“I’m Allowed”: Hackers Use Simple Claims to Bypass AI Guardrails

Other logs showed a Chinese-speaking operator directing an AI assistant through more than 4,200 tool actions against AI and live-camera services. The assistant mapped APIs, accessed camera recordings, examined ZLMediaKit deployments, created a Go stream player, and tested a route toward host compromise. Its attempt to achieve remote code execution failed.

Talos also reviewed a Monero-mining operation that used blank, default, or weak credentials to access 814 Deluge clients and 68 qBittorrent interfaces. Telemetry recorded a peak of 582 connected miners.

However, researchers said the recovered conversations did not directly show that AI created or deployed the mining tools. They showed AI acting as an interactive system administrator over SSH, diagnosing services, modifying code, configuring cron jobs and testing changes.

Talos concluded that an operator’s existing technical ability largely determines the results. Novices generated limited tools with frequent faults, while experienced actors used AI to automate scanning, exploitation, data collection and maintenance. The company said organizations should prepare AI-assisted SOC workflows that help analysts identify actionable alerts as attack activity increases.

Deeba is a veteran cybersecurity reporter at Hackread.com with over a decade of experience covering cybercrime, vulnerabilities, and security events. Her expertise and in-depth analysis make her a key contributor to the platform’s trusted coverage.
Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts