Get started
ArchitectureMulti-agentOrchestrationEvidence

How to give an agent hundreds of tools without sending hundreds of tools

Give one agent every tool your platform has and you hit a wall: send everything and blow the provider limit, truncate and quietly lie, or go narrow and blind the orchestrator. The fix is to unbundle awareness from authority. Here are the measured numbers.

Adam Cowles2 October 20265 min read
A bright central node bound to three hexagonal tools inside a vast faint field of distant points, with a thin chain thread trailing off, cyan and violet on near-black

Give one agent access to everything your platform can do and you hit a wall that every team building multi-agent systems eventually hits. We hit it hard enough to name it, and then we built our way out: two Quox features, the Mirror and QuoxHat, that between them let an agent be aware of hundreds of tools while carrying almost none. This is how they work and what they measure on our own running system.

The trilemma

An agent is only as useful as the tools it can reach. So you add tools. Then you add more. Past a few dozen, three bad options are all you have.

The first is to send every tool on every turn. The cost of that scales with your catalogue, not with the task, and it runs straight into provider limits. Ours did: one provider caps a single request at 128 tool definitions, and the day a wide agent crossed that line its governed failover to that provider died silently, because the request carried more tools than the limit allowed and the whole attempt was rejected.

The second is to truncate. Cap the list, drop the rest. Now the platform quietly lies about what it can do, and which capabilities vanish has nothing to do with what the user actually asked for. A band-aid that leaves the agent deficient in a way no one can see.

The third is to make every agent narrow and route between them. This is the right instinct, and it is where most teams get stuck, because delegation dies without awareness. An orchestrator cannot hand work to a specialist it cannot see. Make the specialists narrow enough to be cheap and you blind the one agent whose whole job is knowing who to call.

The thing that was actually wrong

The three options feel like a genuine trade-off because they share a hidden assumption: that to know about a tool, an agent has to carry it. Awareness and authority were the same thing. Unbundle them and the trade-off disappears. That unbundling is exactly what the Mirror and QuoxHat are.

The Mirror is awareness. It is the platform's model of itself, every agent, tool, plugin and capability, as names and purposes and counts, never as schemas. A human reads it on a page. An agent reads it through a single tool call. It is the same index, and it is cheap, because a name is not a schema.

QuoxHat is authority. A hat is a small, named, bounded profile of tools, resolved on the server for every single turn into a frozen manifest with a cryptographic digest. The agent carries the manifest, nothing else. If it needs more depth, it loads a bundle, which is one governed, audited step and takes effect on the next turn. If it needs a different speciality, it delegates, and the specialist's own turn wears its own focused hat. If it needs a privileged mode, a human approves the switch; it is never automatic.

One sentence holds the whole design. The agent is aware of everything through the Mirror, carries almost nothing but its hat, reaches anything through a governed step, and every step is witnessed.

The numbers, measured not asserted

On our own running system these are the collector's own estimates of schema weight per turn, for the same agent on the same organisation. An unbounded turn carries about 136 tool schemas, roughly 23,790 tokens, before the user has said anything. Under the default hat the same agent carries 10 tools, about 2,545 tokens. When it delegates to a specialist, the specialist's turn carries 3 tools, about 362 tokens. On our main working organisation, where the catalogue is larger, an unbounded turn measured about 28,473 tokens against the same ten-tool default.

The cost per turn stopped scaling with the catalogue. That is the claim, and it is the only efficiency claim we will make. Whether a tighter tool set also makes the answers better is a reasonable guess, and we have not measured it, so we are not going to tell you it does.

Two-level delegation, with a receipt

The part we are most pleased with is the part that usually breaks. A request arrives at the orchestrator, wearing a small general hat. It recognises the work belongs to an infrastructure specialist and delegates. The specialist's turn resolves its own hat automatically, from the agent it is, and wears a focused infrastructure profile. Neither agent ever held the other's tools. This is the kind of governed hand-off the QuoxFlow workflow engine is built to run.

And because the whole point of this platform is that nothing happens without evidence, that delegated turn leaves a receipt. The manifest the specialist wore, its digest, which hat it was and that the hat was chosen automatically, all of it is witnessed in a tamper-evident chain. You can ask, after the fact and without any special access, exactly what any agent could do at the moment it did it.

What we are not claiming

Enforcement runs live on our own working organisation and a scratch test organisation; it is not switched on across a fleet. One provider's mid-walk failover under an enforced manifest is covered by tests but has not been driven against the live provider. We tell you this here for the same reason the design witnesses every turn: a claim without its evidence is just a nicer way of guessing.

Why we think this generalises

The trilemma is not specific to us. Any system that gives agents real capability and expects them to delegate will meet it. Separating awareness from authority is a small idea with a large consequence, and we suspect it is the missing piece in a lot of multi-agent architectures that currently choose one of the three bad options and live with the cost.