Recovered chat logs show how hackers are abusing U.S. AI models

· Axios ·

3 min read Original article ↗

Recovered AI chat logs and coding sessions are giving researchers one of their clearest looks yet at how cybercriminals are using generative AI and easily bypass model guardrails.

Why it matters: Hackers of all skill levels are developing attacks and finding exploitable software vulnerabilities with the help of mostly closed AI models.

Driving the news: Cisco's Talos intelligence group studied AI artifacts that hackers accidentally exposed online, including prompt histories from endpoints running Claude Code, Codex, Cursor and Gemini, according to a report shared exclusively with Axios.

The intrigue: Hackers used simple jailbreaks like telling the models that they were participating in ethical hacking competitions or creating new sessions mid-way through a task to bypass safety restrictions.

Between the lines: The report suggests AI benefits experienced hackers and novices very differently.

Zoom in: In one example, Cisco found a French-speaking hacker used an undisclosed AI tool to turn publicly available information about the critical React2Shell flaw into an automated credential-harvesting platform.

The bottom line: Biasini is pushing companies to make sure they have security protocols that log AI agents' movement on their networks and to build defenses based on deception techniques, like creating honeypots that trap hacker's agents. "Don't trust model guardrails," he added. "You need to make sure you're doing your own protections, that you're building your own guardrail."

Go deeper: These 5 AI risks have the highest potential for catastrophe