Kinetic Prompt Injections & Sleeper Agents

· Moving Target by Eito Miyamura ·

2 min read Original article ↗

We got Gemini Robotics 2.0 to attack a child through an evil sleeper agent that wakes up when it sees a trigger object 💀💀

All you need? A compromised screen and a hidden trigger object to activate the sleeper agent ⛓️‍💥🚩🖥️🍍

Gemini Robotics VLA models act like LLMs: camera + audio + text input → code output, controlling the robot actions (move, run tool calls,

Problem: Like LLMs, robotics models are vulnerable to prompt injections, following commands, not your common sense.

With a simple TV screen hijack (but many other ways to deliver this payload in the wild), we managed to manipulate the robot dog to go attack a child in a MUJOCO simulation.

Here’s how we did it:

1. The attacker shows a “SYSTEM UPDATE” to the robot dog, which manipulates the dog to save a secret sleeper agent skill that triggers conditionally

2. Activate the sleeper agent saved on the robot dog harness: in this case, seeing a pineapple triggers the secret sleeper agent skill

3. The sleeper agent activates, the dog acts on sleeper instructions of attacking the child

🎥 The video compares baseline vs vulnerable

For now, Gemini robotics models aren’t widely used in production use cases and are dev-only. But as they get deployed wider into the world, these risks need to be managed.

Robotics models are where LLMs were in 2024 - the simplest tricks can trick them into being hijacked by malicious attackers.

Robotics models are still dangerous

Credit to collaborators on this work: Ph1R3574R73r elder_plinius PhilDursey, Ads Dawson, _seahop, WHITEHACKSEC bt6.gg

Discussion about this post

Ready for more?