Least privilege for AI agents: scope the blast radius before anything goes wrong

Reason
AgentsQuoxMindAgentic TeamsMirror & QuoxHatQuoxLensRemember
QuoxMemoryBrain2CompoundingQuoxPlanCodebase MirrorAct
QuoxFlowQuoxEngineQuoxAgentQuoxChatAutonomyRun
EnterpriseOrganisationsPowers & ToolsmithQlusterQuoxBastionInterfaces
QuoxMCPQuoxCLIQuoxTerminalQuox ConsoleQuoxBoxQuoxVoiceGovern
HITL ApprovalsQuox SecurityAgent HonestyQuoxVaultAI GovernanceAgentic AIProve
For AuditorsVerifiable AI OpsLoggingQuoxSEOEU AI ActCompliance SuiteChannels
Matrix RoomsDiscord ProTelegram ProQuoxSignalCoreCommsAll products A-ZBuild
QuoxProofDeveloper KitPlugin SDKBring your tool to QuoxQuoxpertQuoxSkillsShip and sell
Build and sellBrowse MarketplaceDownloadsProtocolsDev Suite
Dev ServersQuoxBuildQuoxSlotsShared Skills + RulesQlarityDev workflow
RepoBrainGripeTriageFixLoopProofLoopDoneEngineAll products A-ZGet started
OverviewArchitectureProtocols
AEEAOCLVOLTWARDReference
GlossaryAPI ReferencePlugin SDKDockerAll products A-Z
The instinct with an AI agent is to ask how clever or careful it is. The more useful question is colder: if this agent were tricked, compromised, or simply wrong, what could it actually do? The answer is not a property of the model. It is the sum of the tools it holds, the data it can read, the actions it is allowed to take, and how long that access lasts. That sum is the blast radius, and least privilege is the discipline of keeping it small.
Least privilege does not make an agent behave. It assumes the agent will, at some point, misbehave, and makes sure the damage is bounded when it does. You scope the access first, and then it matters far less whether the model was fooled.
Access is not one dial. An agent is scoped on four separate axes, and leaving any one of them wide open quietly undoes the other three.
Tools. Give an agent the narrowest set of tools its job requires, and nothing it does not. A summarisation agent that reads a mailbox does not need a tool that sends mail to an arbitrary address. Most agent incidents are not exotic; they are an over-scoped agent doing exactly what it was permitted to do. Every tool an agent holds is a verb an attacker can borrow.
Data. Scope what the agent can read as tightly as what it can do. An agent working for one team should see that team's data, not the whole organisation's, and never another customer's. Data reach is where a small prompt injection turns into a large exfiltration, so the default is the narrowest scope that lets the job happen, not the widest that happens to be convenient.
Actions. Not every action deserves the same freedom. Reading is cheap and reversible; moving money, deleting records, or mailing a customer list is neither. Scope the high-impact, irreversible actions so they stop for a human, while the routine ones flow. Approval is a checkpoint on the few steps that would hurt, not friction on all of them.
Duration. Access should expire. An agent that needs elevated reach for one task should hold it for that task and lose it afterwards, through a time-boxed lease rather than a standing grant. Standing access is a liability sitting idle until the wrong instruction reaches for it months later.
| Scoped on | Over-scoped | Least privilege |
|---|---|---|
| Tools | Every tool the platform offers, just in case | Only the tools this job needs |
| Data | The whole organisation, or every tenant | This team's scope, and no wider |
| Actions | Any action runs the moment the model decides | Irreversible actions stop for a human |
| Duration | Standing access that never expires | A time-boxed lease that lapses after the task |
The appeal of least privilege is that it does not rely on the model being clever, honest, or impossible to trick. The limits are enforced by the platform, outside the model, and injected text or a confused plan cannot argue with them. A tricked agent still cannot reach a tool it was never given, read data outside its scope, take an irreversible action without a human, or use access that has already expired. You stop betting on behaviour and start bounding consequences.
This is the same posture as defending against prompt injection: assume the agent will be fooled, and make sure it cannot do much when it is. Least privilege is where that assumption becomes concrete settings rather than good intentions.
Quox scopes all four axes by default. Agents run with a declared set of tools and a data reach bounded to their organisation and role, high-impact actions pass a policy gate and stop for human approval, and elevated access is granted as a time-boxed lease rather than a standing permission. No single one of these controls is novel; the point is that an agent's blast radius is a deliberate decision rather than an accident waiting to be discovered.
The broader argument, that security is about the access an agent holds rather than the model behaving, runs through the security page. How the policy gate and approvals work in practice is set out on the governance page.