Get started
Security & GovernanceMulti-tenancyIsolationEvidence

Can another tenant read your AI agents' data? Make them prove they can't.

A multi-tenant breach rarely looks like a breach, it looks like a query that forgot a WHERE clause. The six questions that separate real tenant isolation for AI agents from the word "isolated", and why a live, both-directions proof is the only evidence that counts.

Adam Cowles29 September 20266 min read
Two isolated luminous domains separated by an impermeable boundary, cyan and violet over near-black

A multi-tenant breach rarely looks like a breach. There is no ransom note and no dumped database. There is a query that forgot a WHERE org_id = ?, and for some window of time one customer's agent could read another customer's data: a decrypted envelope, a run's evidence, a page of search traffic, a plan. Nobody announces it because nobody sees it. It is a diff, not a headline.

If you are handing an AI agent platform your credentials, your fleet access, your plans and your agents' memory, this is the question underneath all the others: when the platform holds two customers at once, what actually keeps yours from leaking into theirs? Not the marketing word "isolated". The mechanism. And the proof.

Why agents make the tenancy problem sharper

Ordinary SaaS multi-tenancy is already hard. Agents make it harder for three concrete reasons.

They carry cross-cutting context. An agent stitches together memory, past runs, tool results and a live request. A single place where the code reaches for "the org" and finds the wrong one taints everything downstream: the memory it writes, the summary it returns, the next decision it makes.

They act, not just answer. A read leak is bad. An agent that resolves to the wrong tenant can also write, dispatch a tool, or spend a budget against the wrong account. The blast radius is actions, not just rows.

They lean on defaults. The most common real cause is not a clever attack. It is a fallback. A background job with no request context reaches for org_id || 'default', and now every tenant's work is tagged to one shared bucket: the same lineage, the same ledger, the same crystallised "learning".

We have found exactly this shape in our own code and fixed it; it is the boring failure that ships, not the exotic one.

What real isolation looks like

Strong tenant isolation is not one control. It is the same rule applied everywhere, with no exceptions that a request can reach.

  • The org is derived from identity, never from the caller. The tenant a request runs as comes from the authenticated session, not from an org_id in the query string or an X-Org-Id header a client can set. A platform that trusts a client-supplied org id has no tenant boundary; it has a suggestion box.
  • Every read and every write is scoped, including the subqueries. Not just the top-level list. The anomaly check, the sibling-record stitch, the "related items" join, the aggregate stat. One unscoped subquery is a full leak.
  • Credentials are leased per tenant, just in time. An agent holds a scoped, short-lived credential for its own organisation, not a wholesale key. A prompt-injected agent in one tenant cannot reach another tenant's secrets because it never held them.
  • The honest hard parts are named, not hidden. Some things are genuinely shared and cannot be per-tenant without a redesign: a single global hash chain, a background scheduler, a cache. A serious platform tells you which boundaries are per-tenant today and which are on the roadmap. The vendor who claims everything is perfectly isolated is the one who has not looked.

The questions that separate isolation from theatre

Ask these, and listen for whether the answer is a mechanism or an adjective.

  1. Where does the request's tenant come from? If any answer includes a client-supplied header or parameter, stop there.
  2. Show me one read path and one write path with the org scoping in the code, including the subqueries.
  3. What happens to a background job or a scheduled task that has no user in context? What tenant does it run as, and can it mistag another tenant's work?
  4. Which stores are per-tenant and which are shared? Name the shared ones and what stops them leaking.
  5. How would I, as a customer, detect a cross-tenant read if one happened? Is there a record I can check, or do I take your word?
  6. Can you demonstrate it live: authenticated as tenant A, try to read tenant B's data and show it denied, and in the same breath show tenant A's own data still returned?

That last one matters more than the rest combined, and it is the one almost no vendor offers.

Don't accept the diagram. Make them prove it, both directions.

A data-flow diagram proves nothing. A passing internal test proves the code at one moment, not the running system your data will actually live on. The only evidence worth anything is a live attempt, on the real deployment, that checks both directions:

  • tenant A, authenticated, is refused tenant B's data (denied or empty, never the other tenant's rows), and
  • tenant A's own request still works.

The second half is the one people forget. A "fix" that blocks cross-tenant reads by breaking every read for everyone would pass a one-directional test and fail every customer. Real proof asserts the boundary holds and the product still functions.

Then ask for the evidence in a form you can re-check yourself, not a screenshot. A signed, dated record of the check, on a ledger the vendor cannot quietly edit, is the difference between a claim and proof.

How Quox approaches it

Quox is self-hosted, so the first isolation boundary is your own network: there is no shared cloud tenancy to leak across in the first place. Within a deployment, the tenant is locked to the authenticated session, reads and writes are org-scoped down to the subqueries, and agents receive just-in-time, organisation-scoped credential leases rather than wholesale keys. The security posture around all of this is written up at Quox security.

We hold ourselves to the live-proof standard on the same page we ask you to hold us to: /security carries our own witnessed self-audit, including the findings that were still open, and a hostile-auditor prompt you can run against our public code so you do not have to take the audit on faith.

Cross-tenant isolation is exactly the property we test that way: not "trust the boundary", but drive a real cross-org request against the running system and witness the result, on the same append-only ledger documented at the WARD protocol.

If you are evaluating an AI agent platform, take the six questions above into the call. The vendor who can answer them with mechanisms, and prove the boundary live, is the one you can put two customers behind. The one who answers with the word "isolated" has told you the one thing you already knew you wanted, and nothing about whether it is true.