You created a signature to detect and block an agent using a risky tool? Cute. An agent will find 18 different ways to do the same thing in seconds. Oh, and the most inefficient and risky path to accomplish a task? Yeah, the agent takes that path even when nothing malicious is involved. Have fun with those benign alerts.

Detection engineering has worked the same way for years. A new attack technique surfaces. Detection engineers create a new signature, a specific pattern, to spot the attack technique. As they deploy it to security tooling, false positives naturally pop up. They fine-tune the rule to reduce the noise and redeploy. Rinse and repeat.

There’s beauty in the precision when done correctly and horror in the number of false positives when done poorly.

But that precision comes with a significant tradeoff. Precision means you're very confident about one very specific thing. Change one tiny little detail of the attack signature, and the detection becomes as useful as your cousin who never helps clean up after dinner. That’s how attackers slip past the defenses.

Narrow Scope, Narrow Results

I’ll spare you the “agents operate at machine speed” trope. But agents require a different way of thinking. Like a toddler testing a tired parent, agents will keep pushing the boundaries to get what they want. For an agent, the “want” is to complete its assigned task. If an agent is blocked, it will work to find another way to accomplish the task.

Imagine this. You create a detection rule or policy that tells an agent not to use the “rm” tool to delete a file. The agent respects that block (most of the time) and won’t use that tool. But it still needs to delete that file. So, it goes to its Rolodex of tools and finds another way to delete the file.

Surprise! There are many ways to delete a file. You don’t know all of them. I sure don’t. But an agent does. For example, skip deleting the file entirely and instead zero it out. Same intended outcome, different path. An agent has all the time and tokens it needs to find a way to delete that file. That asymmetry doesn’t favor defenders.

# Different ways to delete a file #

# Deletion Calls #
rm file
unlink file
find . -name file -delete
echo file | xargs rm
shred -u file

# Via Runtimes #
perl   -e 'unlink "file"'
python3 -c 'import os; os.remove("file")'
ruby   -e 'File.delete("file")'
php    -r 'unlink("file");'
node   -e 'require("fs").unlinkSync("file")'
tclsh  <<< 'file delete file'

# Empty Contents to Null #
truncate -s 0 file
cat /dev/null > file
cp /dev/null file
dd if=/dev/null of=file
printf '' > file
echo -n "" > file
tee file < /dev/null

The Perspective Shift

What is 8.5x11 inches? What about 4x6 inches? 27x40 inches? Spoiler alert: they’re a US letter-size sheet of paper, a printed photo (remember those?), and a movie poster. Go up one level, and they’re printed paper. Go up another level, and they’re all rectangles.

When children are learning shapes, we don’t encourage them to memorize every possible rectangle dimension. We teach them the general rules that define a rectangle. It’s a four-sided shape, with four right angles, and opposite sides equal in length and parallel to each other.

This allows you to recognize an object by its overall shape, not its specific dimensions. That’s why anyone over the age of five can reliably recognize a rectangle by its shape.

Start with Threat Shapes, Then Go Deeper

The deterministic approach requires you to define specific parameters. Specific tool calls or arguments. They have their place in detections when you need the highest fidelity possible. But the non-deterministic and creative nature of agents means we have to open our aperture wider. If you only start with the detailed view, you will absolutely miss things. A lot of things.

In our analysis of System One models, which function as a quick gut-check for a question, the results shocked me. Deterministic pattern matching caught a measly 1.5% of risky actions. With an out-of-the-box System One model, we caught 96%. That includes all sorts of defense evasion techniques we tested that typically bypass detections. You don’t have to be a math major to see that difference.

This is possible because a System One model evaluates the given input and returns a confidence score. Instead of looking up a table of all possible rectangle measurements, it knows what a rectangle is and looks for the general shape.

Back to our delete example. A System One model takes the tool command, determines what it is, assesses the potential outcome of that tool call, and returns a likelihood that it results in a delete operation.

On its own, it won’t give you the right final answer. It doesn’t need to. Think of this as a filtering tool. Take the thousands of actions an agent takes and spot the 150 that could be risky, whatever that definition is for you. It looks like this:

  1. Identify Threat-Shaped Activity: Surface risky activities

  2. Enrich with Context: Add context to understand what is happening. What led up to that action? What’s normal for the human or agent? What does this action lead to? This is where the risky-but-harmless activity gets sorted out before it becomes another false positive.

  3. Determine the Threat: Threat-shape + context enables an automated decision on whether activity is malicious, benign, or needs escalation to human judgment.

The power is in the full system. Like a three-legged stool, missing just one leg creates an imbalance that only favors gravity.

Agentic Threat Shapes

It’s time to lift our heads above the canopy of the signature forest. Deterministic patterns have a place in agent detection and response. We won’t abandon them completely. But we need to rethink what detection means for agents. That’s why at Evoke, we monitor agent activity looking for threat-shaped patterns as an initial filter.

Detections start there and then run through our context engine to understand what an agent is doing and why. That includes our 4M approach to identify rogue actions, because bad intent is only one way an agent goes rogue.

An agent's dedication to completing a task shifts the detection dynamics. There’s no longer lag time between technique and control bypass. It’s happening in real time. Stop looking for specific threat measurements and start looking for threat shapes.

If you need a partner to help find those threat shapes in agents, let’s chat.

Reply

Avatar

or to participate