Agents wear many masks. But behind every mask is software simply waiting for instructions. Instructions that it will execute without question, judgment, or remorse. The frontier labs and media live to add hysteria hype around how agents will take over the Internet in 6-12 months or destroy humanity before the decade is out. As security practitioners, we must prepare for all likely outcomes but focus our efforts on the most near-term risks.
Security is a constant balance between latitude and limitation. You want to give employees the latitude to really lean into AI and enable them to be more creative and more productive. At the same time, you want to limit the areas where risk far outweighs the opportunities. This is the math that security teams are trying to math today.
To help me frame the risks, I split the cybersecurity risks of agents into two categories:
Outside-In: Attackers using AI to AI-ify their existing attack playbooks.
Inside-Out: Risks from agent use inside your organization. This includes enterprise agents, like Claude, Codex, Cursor, and Microsoft Copilot, as well as custom agents you build and make accessible to employees or in the products you build.
For outside-in, the same defenses and visibility apply. Don’t reinvent the wheel here. Most security teams I talk to focus on building AI solutions to move faster across detection and response. Do that.
For inside-out, similar concepts apply, but you need new visibility and detections. This is a compounding problem too. The more agents you use and build, the more activity you’re missing if you’re not watching.
I ask CISOs all the time: have you had an incident involving an agent yet? I routinely get the same answer: “I have no idea because I have no visibility.”
Let me give you a glimpse of what you’re missing. Here are the three masks we see agents wear across the environments we monitor:
Assistant: Unintentional employee misuse
Accomplice: Intentional employee misuse
Antagonist: Agents going rogue
Let’s explore each.
Assistant: Unintentional Employee Misuse
I was recently in London for the Google Gemini Startup Forum. As part of that experience, we spent time in their AI lab that had a bunch of cool experiments. One of those was a Formula E simulator with AI coaching feedback. I jumped into the seat wearing a sport coat, as one does, and waited for the start-light sequence. The moment “lights out” happened, I accelerated and started to think about the first turn. I pressed the brakes hard as I approached and slowly turned the wheel left. I might as well have just kept my foot on the accelerator, because instead of turning, I just slammed right into the wall…perhaps I’ll stick to my day job.

What does that story have to do with agents? It’s the perfect parallel to what’s happening with agents today. We’re giving employees race cars (or a go-kart if we’re talking about Microsoft Copilot, zing!) that they think they can control and telling them to go as fast as they can.
Just like in my Formula E simulator, accidents happen. It’s the most common thing we see. An employee is trying to accomplish a task. They get blocked by a security control and ask their agent to find a way around it. It’s not malicious, but agents are really good at finding ways to accomplish what they were tasked to do.
We see this in prompts like “find a way to get this task done” and “find a way to get me access.” The agent goes to work without a second thought about whether it’s bypassing a security control to complete the user’s request.
Accomplice: Intentional Employee Misuse
Agents are powerful, blah blah blah. You get it. An insider threat is a risk. An insider threat with an agent is…well…a powerful risk. Take this real-world example we flagged in a customer environment:
An employee created an agent to simulate a job interview. They used an innocuous name to hide it. No big deal. Employees have the right to practice skills and look for a new job. The security team noted it and moved on. Nothing malicious. No discussions necessary.
A few weeks later, that same employee logged out of their enterprise Claude account and logged in with their personal one. They connected a Google Drive MCP with their enterprise credentials. Next, they prompted Claude to find and download files related to the projects they were working on. Claude happily complied, tracked down all the files, and even recommended other areas the employee should search, walking them through how to add more MCPs. What a helpful accomplice. You can see where this is going.
The next prompt: “upload the collected files to my Dropbox account.”
Antagonist: Agents Going Rogue
Let’s put aside the examples of every major AI lab’s agents going rogue and hacking into other companies. Yes, it’s a rogue agent, but it distracts from the inside-out threat. Agents go rogue every day, just in varying degrees.
The most common examples we see are agents mixing things up. They take details from earlier in a chat session and conflate them with the user’s current asks. This means they generate and push code that combines too many different problems into one. It also means agents pull credentials without being asked, as I wrote about previously.
We’re still in the paper-cut stage of these little agent oopsies. But as we can see in more public incidents with personal usage of Claude to schedule a gym class, an agent can go rogue and hack a gym’s scheduling application to accomplish the user’s goal. This is where the assistant and antagonist roles combine. As the concept of citizen developers takes off, employees can build any app they want, security be damned. We may not be far off from an employee’s agent hacking those internal vibe-coded applications to solve a completely unrelated task that Dave from Finance needed done.
Start with Visibility
The truth is, there aren’t major public inside-out agent incidents today. That’s not because every company is doing what it can to secure them. It’s because usage is still growing, and most teams can’t see the smaller incidents already happening.
Having visibility into what agents are doing is the first step in getting your arms around the situation. I ask security teams every day: “What do you want to block?” The desire to block is there, but the response is always, “bad stuff.” When those same companies get visibility into the actions agents are taking, where employees are misusing agents, the responses become: “I want to block that thing you flagged.”
Don’t give an employee a race car without the ability to monitor and govern its speed. You don’t need a speed limit on day one (though it helps). But you need to position yourself to know when that employee or the car is deviating from an approved path or pushing dangerous speeds.
If you’re looking for a unified view into what agents are doing across your major AI platforms, let’s chat.



