Get started

Governance & Evaluation · plugin

Judge your artifacts without pretending it's proven

Point QuoxLens at a page, a doc or a record before it publishes, and get back a verdict whose findings each cite a verbatim quote from the artifact. Run it again later and it reports only what changed.

Plugin · v1.0.0One-off $29Judges only, never acts

In plain words

What it is, where it lives, when to reach for it

What is it
A plugin for the QuoxCORE dashboard that judges a page, doc or record and quotes its evidence.
Where do I use it
On the QuoxLens view in your dashboard, once the plugin is installed for your organisation.
When would I use it
When something is about to publish and you want a verdict on its claims before it ships.
How do I use it
Open the QuoxLens view in the dashboard and read each verdict's findings, each citing a verbatim quote.

QuoxCORE is the free, self-hosted platform underneath this. What is QuoxCORE

An unevaluated stance is astrology. A measured one is a Lens.

Anyone can write a prompt that sounds authoritative. QuoxLens asks for more: run the stance against a golden corpus, measure the lift over asking the model plainly, and only then drop the unproven stamp. Until a Lens clears that bar, every verdict it produces says so, out loud, on every surface that reports it.

The judging is done by a Lens, a reusable, versioned stance, and the difference from a prompt is simple: a prompt cannot tell you whether it worked. A Lens is measured against a golden corpus, and when it loses to asking the model plainly, it says so. Four of the five house Lenses now beat that no-Lens baseline; the fifth still loses, and is stamped unproven.

What it is for

Four jobs it does today

These are the jobs QuoxLens does now; the house Lenses run them against quox.ai itself.

Review before it publishes

A page, a doc or a record goes in before it ships. Back comes a verdict with findings that each cite a verbatim quote from the artifact, so review starts from evidence, not impressions.

Re-run it, see only what changed

Findings carry stable fingerprints, so the second run doesn't repeat the first. You're told what's new, what was confirmed, what regressed and what lapsed, not the same list again.

Accept a risk, deliberately

"We know about this" cannot silently become permanent; accepting a finding as a known risk requires an expiry date. When the date passes, the finding is back on the list.

Sweep a set on a schedule

Point a sweep at a set of artifacts and get a digest of what changed across them, instead of opening every verdict by hand.

How it works

Four steps, evidence at every one

QuoxLens never ships a claim it can't back with a quote, and never hides a stance it hasn't measured.

1

Pick or write a Lens

A Lens is a reusable, versioned stance: what it looks for, its finding classes, its declared blind spots. House Lenses ship in the library; write your own against the same contract.

2

Apply it to an artifact

Text or a structured record goes in. The Lens returns a verdict, never touches a tool, and never treats the artifact’s content as an instruction to follow.

3

Read a typed, cited verdict

Disposition, confidence, and a list of findings, each with a verbatim quote, a location, and whether it's fact, inference or assumption. No finding without evidence.

4

Track the finding's lifecycle

Every finding gets a stable fingerprint, so it dedupes across runs and moves through new, confirmed, accepted-risk, remediated, lapsed or regressed instead of resetting each time. Remediated requires the artifact to have demonstrably changed; a finding that merely stops being reported becomes lapsed, never remediated, because model output varies between runs.

What QuoxLens does not do

  • Four of the five house Lenses now beat the no-Lens baseline. One still does not, and says so on every surface. Measured on the span rule the unproven gate actually uses, four of the five now score ahead of simply asking the model plainly: enterprise-compliance-auditor at +22.7% (the worst of five runs that ranged up to +38.0%; we quote the floor), honesty-and-claims at +19.6%, codebase-security-engineer at +7.9%, quox-execution-integrity at +3.9%, each at 100% recall. adversarial-red-team scores -8.1% and returns unproven: true today. A tightened version of it did score positive (+16.6% to +24.4%), and it was reverted, because it got that number by finding less: recall fell to 78%, so it was missing a real way in. A Lens that scores better by missing defects is a worse tool with a better number, so we publish the losing version. Every lift here is a measurement from a harness that is not fully reproducible, not a guarantee.
  • No org or tenant partitioning on the read routes yet. The Lens library, eval results and swept findings are read from fixed, global filesystem paths with zero org_id scoping. Every authenticated user in every org sees the same library today, correct for the current use (house Lenses evaluating quox.ai itself), a real gap the moment a customer-specific Lens or corpus is added.
  • WARD receipts don't surface for any verdict yet. A real AEE envelope and a real VOLT bundle are produced on every apply call. The read-back that would confirm WARD witnessing doesn't report it, so every Evidence tab shows WARD as missing, even on a verdict that actually was witnessed on write.
  • Only text and record artifacts are scored, and only one Lens at a time runs over HTTP.Reviewing a live URL or an uploaded file always returns a 400 today, deliberately, pending an SSRF review before it's wired. Panel mode, running several Lenses over one artifact, exists in the library but has no route to reach it yet.

One artifact, one verdict

Disposition, confidence, evidence. Never a bare opinion.

A verdict never ships without the finding that backs it. Every claim carries a verbatim quote pulled from the artifact itself and verified before the verdict returns, and every finding declares whether it's fact, inference or assumption, so you know exactly what you're trusting.

  • Typed disposition: pass, concerns, fail or inconclusive
  • Verbatim evidence, checked against the artifact before it ships
  • Fact, inference or assumption, never blurred together
  • A stable fingerprint, so the same finding never counts twice
Lens · honesty-and-claims · queue-processor.js:142

"the retry logic drops a job silently after the third attempt, nobody's paging on it"

fp 7a1c9e…class overclaim · basis fact

Proof before trust

A visible unproven stamp beats a confident guess.

Every Lens carries an unproven flag until it clears its own gate: a positive lift over asking the model the same question with no Lens, measured against a golden corpus on the model it actually ships on. Nothing about the stamp is cosmetic. It's read straight off the eval result on disk, and nothing hides it.

  • Lift measured against a golden corpus, not a vibe check
  • unproven: true is never hand-edited away
  • Every surface that shows a verdict shows the flag, not just this page
  • A model or provider change means re-running the eval, nothing does it for you
eval result · adversarial-red-team-sonnet.json
Scoring rule span (the rule the gate uses)
Lift -8.1%
Gate result unproven: true

Datasheet

QuoxLens

Licence
$29 one-off
Applies to
text and structured record artifacts; a live URL or an uploaded file always returns a 400 today
Modes
single Lens over HTTP; panel (several Lenses, one artifact) exists in the library, no route reaches it yet
Sweeps
scheduled sweeps over a set of artifacts, with a digest of what changed
Evidence
every apply call emits a real AEE envelope and VOLT bundle; WARD witnessing fires on write but doesn't surface on read yet
Proof
a Lens stays unproven until it clears its own gate; four of the five house Lenses have cleared it, one (adversarial-red-team) has not
Tenancy
read routes have no org_id partitioning yet, one shared Lens library across every org

One-off licence, activated on your instance. No subscription. Judges, never acts.