Our entire pitch is that you should not trust claims, you should check evidence. At some point that has to apply to us.
So we ran an experiment. We took a frontier model from a competing lab, gave it quox.ai and the sitemap, and told it to conduct a cold evaluation as if it had no connection to us. We instructed it to be sceptical. We told it to look for vaporware, inflated numbers, documentation drift, unfinished products, weak moats, excessive branding and governance claims unsupported by evidence. We told it not to let us define our own competitors, and not to soften criticism. We explicitly gave it permission to conclude that this is, in its words, another enormous AI-agent website.
This post is what came back, with the criticism left in. The full prompt is at the bottom so you can reproduce the experiment, on us or on anyone else.
What it said first
The reviewer opened the way most people open an AI platform website in 2026.
When I first opened the site, I rolled my eyes.
Fair. The estate is large. The product names are many. The acronyms are four deep. A cold reader arriving at yet another agent platform has every reason to assume the worst, and we asked this one to.
Where it did not change its mind
An adversarial review you publish is only worth reading if the negative findings survive. These did, and we think they are correct.
It found the named specialist agents commodity. It found the chat interface table stakes. It found that the workflow canvas overlaps with established automation tools. It called the surface area of the platform alarmingly broad for the size of the company, flagged self-hosting as adoption friction, and warned that our control-layer model could prove heavier than some teams want. Its harshest score was commercial clarity, and its sharpest observation was that the marketing sometimes makes serious engineering look like the generic agent-wrapper products it sits next to.
We are not going to argue with most of that. Some of it is already changing. Some of it may simply be the cost of the architecture we have chosen.
Where it changed its mind
The reviewer reported the places where digging altered its assessment, which was the part we most wanted to see.
The four protocols were the first turn. It went in expecting proprietary naming for commonplace concepts and came out treating the envelope, control-layer, evidence and witnessing work as genuine protocol design, with external artefacts it could verify rather than take on faith.
The second turn was coherence. It followed the same primitives, identity, organisational scope, authority, credentials, approval, evidence, from the terminal to the workflow engine to the MCP surface to governance, and concluded the estate is one architecture expressed in several products rather than several products dressed as a platform.
The third turn was the thesis itself. Its summary of what we are building was one we would happily have written ourselves:
They recognized that the blocker for enterprise AI isn't capability, it's liability.
Anyone can make an agent capable of restarting a database. The hard problem is being able to allow that capability into production while proving who initiated it, why it was permitted, what it touched, who approved it, and whether the record can be independently verified afterwards. That is the problem this company exists to work on.
Its closing note, written for a hypothetical investor who asked whether this site was worth their time, was three words: take the meeting.
Where the reviewer itself overreached
Publishing a flattering review without correcting its errors would defeat the point of the exercise, so here are the places we think our examiner got carried away.
It described the evidence and witnessing layer in terms that drift toward zero-knowledge proofs. That is technically sloppy. Cryptographic witnessing and content-minimised receipts are strong properties, and they are not zero-knowledge proofs, and we do not claim they are.
It labelled parts of the platform production-ready from public evidence alone. That is more confidence than a cold review can justify, and more than our own published maturity tables claim. Our documentation marks components Stable, Beta, Experimental and Not built, and we run automated gates that fail our own builds when a page claims more maturity than the status documentation supports. The honest labels are the ones on the site, not the reviewer's upgrade of them.
And its protocol analysis concentrated on one of the four specifications when the external work is broader. A cold reviewer cannot know everything. That is rather the point of publishing the method instead of just the verdict.
The obvious objection
One run of one model is not proof of anything. Language models can flatter. A review conducted this week could read differently next week.
All true, and it is why the deliverable of this post is not the verdict, it is the prompt. We are keeping it as a standing test. When the estate changes, we can run it again and see whether the findings move. You can run it today, against us or against any platform making similar claims, and see what an instructed sceptic finds when nobody is steering it.
That is the same posture the product takes: do not trust the narrator, check the record.
The prompt
Condensed for length, functionally identical. Paste it into any capable model with browsing.
Independently evaluate this company and product ecosystem: https://quox.ai
and https://quox.ai/sitemap.xml. Treat this as a completely cold evaluation:
assume I have no connection to the company, do not try to please me, and do
not infer what answer I want. Browse deeply via the sitemap, not just the
homepage: products, docs, developer material, governance and compliance
pages, blog, protocol specifications, and any external evidence you can
verify independently. Then research the current competitive landscape
yourself and decide what category this actually belongs in; do not let the
company define its competitors for you. Be sceptical: look for vaporware,
inflated numbers, inconsistent claims, documentation drift, unfinished
products, weak moats, architectural over-complexity, excessive branding,
unclear target customers and governance claims unsupported by evidence. For
every important claim, distinguish claimed vs demonstrated vs externally
verified. Then give me: what it is in five sentences; closest real
competitors and a capability matrix; what is genuinely differentiated and
what is commodity; whether the architecture is coherent or feature
accumulation; a technical and strategic assessment of its protocols; your
own product maturity estimates, not the company's labels; who would pay
today; the moat, if any; the ten most serious red flags; anything that
changed your opinion while digging; a 0 to 10 scorecard with reasons; and a
final verdict written as a confidential note to a technically sophisticated
VC partner who asked whether this is actually interesting or just another
enormous AI-agent website. Do not soften criticism. Do not manufacture
praise. Cite URLs and distinguish evidence from inference.
If you run it and find something we should hear about, tell us. Finding out is the whole idea.