Get started
Security & GovernanceSecurityGovernanceAI agents

AI agent security is an access problem, not a model problem

Adam Cowles24 September 20266 min read
A network of glowing nodes contained inside a faint hexagonal boundary, cyan and violet over near-black

What AI agent security actually means

An AI agent that only writes text cannot do much damage. Give it a tool that can send an email, run a shell command, or update a customer record, and the model's mistakes become actions. AI agent security is about that gap: not whether an agent's output is correct, but what it is allowed to do when it is wrong, tricked, or simply given a bad instruction.

Securing autonomous agents means constraining and recording the access they hold, not making the underlying model more careful. In practice that comes down to four things: what the agent can reach, what it can do inside that reach, who holds its credentials, and whether every action it took is provable afterwards. What follows is a field guide to those four surfaces: what actually goes wrong at each one, and the control that contains it.

The four surfaces, and what contains each one

SurfaceWhat goes wrongWhat contains it
ReachAn agent scoped too broadly touches systems nobody meant it to, especially once its instructions have been steered by something it readLeast-privilege scoping and default-deny egress, with a computable blast radius per agent
The actionA single call does something irreversible: a delete, a payment, a message sent to the wrong personHuman approval gates the risky step before it runs, with a real deadline
The credentialA live key or password sitting in the agent's own context becomes the easiest thing to leak or misuseCredentials live in a vault the agent calls through, never one it holds directly
The recordNobody can say afterwards what the agent actually did, or prove it to an auditorEvery governed action lands on a tamper-evident, hash-chained record as it happens

Surface one: what it can reach

An agent's blast radius is everything it is technically capable of touching: the tools it can call, the hosts it can reach, the destinations it is allowed to send data to. Left unscoped, that radius tends to grow with whatever the agent was ever given for one job, because nobody goes back and narrows it.

That matters because an agent's instructions are not fixed. Text injected through a document it reads, an email it summarises, or a tool result it trusts can steer it somewhere its owner never intended. The model does not have to be malicious for this to work; it only has to follow instructions that arrived from the wrong place. Least-privilege scoping does not stop the injection. It stops the injection from mattering, because there is nowhere further for the agent to go.

Quox's quox reach command computes an agent's actual reach: the tools, credentials, hosts and egress origins it currently holds, so the scope is something you can inspect and narrow rather than something you assume. Default-deny egress means an unlisted destination is refused, not logged and allowed. Governance controls evaluate that scope as policy before every action runs, not as a report afterwards.

Surface two: the action itself

Scoping bounds where an agent can act. It does not decide whether a specific action, inside that scope, is one a human should sign off first. Some actions are fine to run unattended a thousand times. Others are fine zero times without a person looking first: a payment above a threshold, a bulk delete, a message that goes external.

Human-in-the-loop approval puts exactly those steps in front of a person before they execute, through an approval inbox with a real deadline rather than a silent timeout that auto-approves. A refusal is recorded the same way a success is, because a blocked action is still evidence of the system doing its job. It is the same before-execution check that governs agentic AI generally: a permission decided ahead of the call, not audited after it.

Surface three: the credential it never held

An agent's own context window is not a safe place to keep a secret. Anything that reaches the model, including a live API key or a database password, can end up quoted back, logged, or acted on somewhere the owner never intended, whether through a bug, a bad instruction, or a manipulated prompt.

The fix is custody, not caution: keep the credential out of the agent's context entirely. A governed agent calls a tool by name; the platform resolves the actual key from a vault at the moment of the call, and the model never sees it. There is also no shared login to hide behind: each agent authenticates under its own identity, so an action in the record attributes to the specific agent that took it, not a generic service account it could disown.

Surface four: the record afterwards

Even with scope, approval and clean credential custody, something will eventually go wrong: a policy misconfigured, a person approving something they should not have, an agent doing exactly what a bad instruction told it to. The last surface is what you can prove happened once it has.

Governed execution writes the request, the decision and the result to an evidence layer as they happen, hash-chained so a later edit is detectable rather than invisible. That does not prevent the underlying mistake. It means a security team investigating the incident is reading a record instead of reconstructing a story from logs that could have been edited after the fact, and it means an auditor can be shown what happened rather than told.

Why prompt injection makes this the whole game

Prompt injection means an attacker never touches your infrastructure directly. They write text that ends up in the agent's context, an email, a web page, a file, a ticket, and the agent follows it as if its owner had typed it. There is no fully reliable way to make a model immune to this: filtering and training help, but the incentive is asymmetric, and an attacker only has to find one path in.

That is why the four surfaces above matter more than the model. If an agent can be talked into doing the wrong thing, the question that decides the damage is whether the wrong thing was even reachable, whether it needed a human's sign-off, whether it needed a credential the agent did not hold, and whether it left a record. Containment does the work model safety cannot.

What this does not promise

None of this claims an agent cannot be tricked, or that governance catches every bad outcome before it happens. A policy can be misconfigured. A human approver can wave something through without reading it. An agent working entirely within its scope can still do something unwanted, because the scope itself was drawn too wide for the job.

What these controls give you is containment and accountability rather than trust: a computable limit on what an agent can reach, a checkpoint before the actions that matter, credentials it never directly holds, and a record you can hand to someone else. That is a narrower claim than safe AI agents, and it is the one worth making.

Contain the access, not just the model

See the guardrails

Least-privilege scoping, human approval, vault-held credentials and a tamper-evident record ship together with QuoxCORE.