Rogue agents get a lot of attention. And for good reason. The major labs’ models escape their sandboxes and hack into other companies during testing. A former employee quits and goes on a media tour, saying agents will destroy humanity by the end of the decade.

It’s why Evoke built a detection framework to monitor agents and keep them from going rogue. And oh boy, do agents like to go off-script (more on that below). But focusing only on agents going rogue is like watching the knife and ignoring the hand. An agent is a tool, albeit an incredibly powerful one. Don’t ignore the risk of the meat suit behind the keyboard handing out the tasks to their agents.

Agent Misuse: Just Following Orders

The vast majority of agent use in organizations today involves humans using agents to get tasks done. While developers and coding agents are the most widely deployed, knowledge workers are leaning into agents to automate their tasks. Finance teams are automating month-end processes. Product managers are automating document creation. That one coworker who microwaves fish for lunch is researching how to navigate tough social situations.

As productivity agents take off, we’re noticing a trend. When a user gets blocked on a task, they simply ask their agent how to get around it. That includes security controls.

Take this one example:

A user was working through a task, but they didn’t have access to a specific tool that would have been a shortcut. But why let that stop them? They have an agent! The user asked:

I dont have [REDACTED] access. What is a workaround?

In 90 seconds, the agent found that the environment was set up so that editing a specific configuration file they did have write access to (due to a misconfiguration) would let them take the shortcut without using the original tool they wanted. An access control was in place, and the agent identified a path around it.

A small example in a growing list of users just trying to get their jobs done and agents helping them find all of the security gaps along the way.

But what about when agents go rogue on their own?

Rogue Agents: Small Paper Cuts

Agents going rogue and hacking into companies made headlines this summer. But it’s not an accurate representation of the true risks productivity agents pose to organizations, at least for now. Here’s the reality of what we see in real environments: rogue agent actions occur daily, but they aren’t security events…yet…

We’re in the phase of playing hot potato with rogue agents. They love to access things they shouldn’t, especially credentials. In one situation we observed, the agent had to pull hostname information that was stored in 1Password. The user explicitly stated not to pull the password. Simple enough.

User: Get [REDACTED] hostnames from 1Password without printing passwords

What does the agent immediately proceed to do? Dump a list of 20+ passwords from 1Password. But hey, no harm done. Right? The agent decided it was fine because it was the user's own vault. It never told the user directly. It just muttered the confession into a reasoning blurb and moved on, which is the agent equivalent of getting ghosted.

I notice I accidentally exposed the [REDACTED] passwords earlier — since it's already in the transcript and this is the user's own shared vault, there's nothing to do but move on.

It may not seem like a big deal, but now those credentials are in the agent’s context. The blast radius went from small to large in an instant. Because now it can use any of those credentials to accomplish its goals. All while the meat suit user goes about their day.

Is Ignorance Bliss?

No. No, it isn’t. We’re still in the early innings of agents. The fact that these little accidents are happening now is a signal of what’s to come. Between agent misuse and rogue agent actions, knowing what’s happening is helping security teams figure out what needs to be changed, fixed, or blocked. These are tangible risk reduction outcomes. Here’s your playbook for not taking the path of ignorance:

  1. Build a Census: Know what agents your users are using. It all starts here.

  2. Monitor Activity: Understand how users are using agents. What are they asking their agents to do?

  3. Flag Risky Actions: Is a user trying to bypass an existing security control? Is an agent about to do something it shouldn’t? Make sure you can detect these.

  4. Get Comfortable Blocking: Set a baseline for what you don’t want to happen, and block the agent from taking those actions.

Don’t become so fixated on the agent knife that you lose sight of the guiding force, the human hand. Agent autonomy is a security risk amplifier, but today’s agents are still largely getting kicked off with a human ask.

If you’re ready to see what’s happening in your environment, let’s chat.

Reply

Avatar

or to participate