The industry has answered this in two ways, and both are necessary. Identity platforms treat agents as non-human identities, bringing them ownership, lifecycle and entitlement review. Inline control products sit in the request path, governing what may reach a model. Both are preventive. Both are worth having. Neither answers the question a security team asks at 2am:
What did this agent actually do, and was that normal for it?
That is not a governance question. It is a detection question, and it needs the same machinery as every other one: telemetry, a schema, a baseline, a framework to classify against, and a way to correlate the result with everything else happening on that entity.
Prevention cannot close this alone
An agent operating exactly within its granted entitlements can still be the vehicle for an incident.
A support agent has legitimate read access to a ticketing system and a document store. It reads a ticket containing instructions addressed to it, follows them, retrieves documents it has never touched before, and summarises them into an outbound reply. Every action was authorised. No entitlement was exceeded.
What actually happened is visible only as a pattern: an unusual retrieval set, for this agent, shortly after an input carrying injection markers, ending in a message leaving. That pattern is a detection, and it exists only if you kept the events and can join them.
The industry collects these feeds in exactly the wrong order
Agent behaviour does not arrive as one log. It arrives as a handful of distinct feeds, and each answers a different question. Prompt and completion text tell you what a model was asked and what it said. Tool and plugin invocation records tell you what the agent then did. Retrieval events carry the document ids and tenant labels a cross-tenant leak would show up in.
What you can detect follows directly from which of these you hold. With prompt, response and tool-call telemetry, indirect prompt injection, tool-invocation abuse, guardrail bypass and retrieval poisoning are all in reach.
Most environments start with prompt logging and little else. It arrives first because it is what the model platform gives you by default and what the compliance conversation asks for. Tool- invocation logging arrives last, if at all, because it has no single obvious owner — it comes from agent-framework callbacks, an MCP server audit log or OpenTelemetry GenAI traces, each of which needs someone to have instrumented the application for it.
That order is precisely backwards. It leaves you with an agent that reasons visibly and acts invisibly — you can read what it was thinking and cannot see what it did.
If you instrument one more thing this quarter, instrument the tool-call log at whatever brokers your agents. It is the only feed that records what an agent did rather than what it was asked.
What you collect then has to normalise into the schema the rest of your telemetry already lands in, and each agent has to resolve to the identity that owns it. That is the difference between an agent read a file and an agent read a file class it has never read, using a credential belonging to an employee on leave, an hour after processing an input containing injection markers. The first is a log line. The second is an incident with a scope and an owner.
One hunt, concretely
The most useful agent detection is about behaviour rather than about attacking a model. p l ug i n_ i nvoca t i on_anoma l y looks for an agent calling a tool it has never called before, or chaining tools in a way that ends with data leaving. The worked example is r e t r i eve_doc followed by s end_ema i l.
Nothing about it is exotic. It scores rarity over the combination of tool, argument shape and calling identity, together with the shape of the chain, against a thirty-day baseline. An agent that has called r e t r i eve_doc many times and s end_ema i l never is not suspicious; the same agent doing both in one chain for the first time is a lead.
It maps to MITRE ATLAS technique AML.T0053, AI Agent Tool Invocation. That technique now sits in a growing family of entries about an agent taking an action rather than a model being attacked — exfiltration via tool invocation (AML.T0086), data destruction via tool invocation (AML.T0101), and tool poisoning (AML.T0110) among them. The catalog is moving toward agent behaviour because that is where the incidents are.
Which is worth pausing on: prompt injection is not phishing, and extracting a system prompt is not credential dumping. Agent attacks have their own preconditions and their own evidence, so ATLAS belongs beside ATT&CK as a first-class namespace rather than folded into it.
Abuse is not compromise
Everything above assumes an attacker subverting an agent. The other failure is quieter: legitimate access used at a volume, or for a purpose, nobody sanctioned. No entitlement is exceeded and nothing is exploited, so a compromise-oriented detection finds nothing — correctly, because there is no compromise to find.
The same distinction predates agents. Samsung's 2023 ChatGPT incident showed the pattern clearly: engineers in the semiconductor division pasted source code and meeting notes into a public model, and within weeks the company restricted generative AI on company-owned devices. There was no adversary in that story and no ATLAS technique to file it under, because ATLAS describes adversaries. What there was instead is a discovery problem and a volume problem.
Agents make that harder in a specific way. When a human pastes source code into a chat window, the actor is an employee with a manager, a working pattern and a baseline someone already monitors. When an agent does the equivalent at machine speed, the actor is non- human, may run continuously, and often has no owner mapped at all. The same failure mode arrives with more volume, less attribution, and nobody watching the clock.
No single feed establishes it. A token spike alone is a busy week. It becomes abuse when it lands in the same short window as a sign-in from a new ASN, a rare tool call, or an egress the agent has no business making. Cost, identity, tool use and data movement, joined on one entity in one window.
That is the argument we have made about phishing. An inbound message is not the incident; treating it as initial access, tied to endpoint and network anomalies inside a short timeframe, is what turns it into one. Agent abuse has the same shape and needs the same join.
What we hold ourselves to
We build with agents ourselves, so the questions we teach customers to ask about their agents are ones we have to answer about ours.
Our agents hold no credentials to customer systems. Access to data runs through a broker, so the credential boundary sits outside the agent. Behavioural scoring is deterministic and statistical rather than model-decided — what is normal for a principal is computed before any language model sees the case, and the model explains the evidence rather than setting the priority or choosing the action. Every model call is traced and inspectable after the fact, and changes are evaluated against a baseline in shadow mode before they reach production.
The hard part is collection
The detection logic for agent misuse is well understood and increasingly well documented. What most environments lack is the telemetry to run it against.
That is the encouraging half of the finding. Collection is a decision an organisation makes rather than a research problem, and it is one a team can make this quarter.
The full technical treatment — feed-by-feed telemetry mapping, ATLAS coverage, and where detection sits beside prevention — is in the companion whitepaper.

.png)



.avif)
.avif)
.avif)
.avif)
.avif)


.avif)
.avif)

.avif)

.png)