Get started
Security & GovernanceComplianceEvidenceCryptography

How an auditor verifies an AI agent compliance evidence pack offline

An AI agent compliance evidence pack has to answer one question on the auditor's own laptop, with no login and no call back to your servers: what did this agent do, was it allowed to, and can you prove it. A worked example of what a pack contains and how it gets verified offline.

Adam Cowles2026-09-28T10:00:00.000Z7 min read
A glowing hash-chain of interlocking sealed links extending across a dark field, cyan and violet over near-black

An AI agent compliance evidence pack is the record an auditor needs to answer one question: what did this agent do, was it allowed to, and can you prove it without trusting the vendor. A good pack answers that on the auditor's own laptop, offline, with no login to your console and no call back to your servers. If verification depends on your dashboard being up and your word being taken, it is not evidence. It is a claim with nice styling.

This piece walks through what such a pack contains and how an auditor checks it, step by step. It assumes you already know why agent governance matters; if you want the ground floor first, read AI compliance, explained, and for the SOC 2 angle specifically, from SOC 2 to AI compliance.

What an evidence pack actually contains

A useful pack is not a folder of screenshots and a CSV. It is a set of signed records, each covering one agent action, that can be verified against each other and against the outside world. The pieces:

  • The action record (AEE envelope). One structured envelope per agent action: what tool was called, with what inputs, by which agent, under which policy decision. This is the payload an auditor reads.
  • The policy decision (AOCL). The record that the action passed the L3 policy gate before it ran, and the L8 verify step after. This is what turns "the agent did X" into "the agent was permitted to do X, and the result was checked".
  • The ledger entry (VOLT). Each record is written into an append-only, SHA-256 hash-chained ledger, signed with Ed25519, and stamped with an RFC 3161 timestamp from an independent time authority. The chain is what makes deletion and reordering detectable.
  • The witness receipt (WARD). A content-free receipt: a cryptographic witness that a given record existed at a given time, carrying no payload. It lets a third party confirm the record is genuine without ever seeing its contents.

The point of four layers is separation. The envelope is readable, the ledger is tamper-evident, the witness is shareable without leaking data, and the policy decision ties the whole thing to a rule that existed before the action, not one written afterwards to fit.

Why a screenshot or a log export is not evidence

A screenshot proves someone saw a screen once. A CSV export from a live database proves what the database said at export time, which is whatever the last writer left there. Both are mutable, neither is independently checkable, and both require the auditor to trust the system that produced them. That is the exact trust an audit exists to remove.

The failure is not theoretical. If an agent deletes a record and the log lives in the same mutable store, the log can be edited to match. A hash-chained ledger breaks on that edit: change one entry and every downstream hash stops matching, and the break is visible to anyone with the verifier and no access to your systems.

The walkthrough: verifying a pack offline

Here is the sequence an auditor runs against a pack that carries all four layers. One caveat up front, because compliance is a high-trust area: the signature layer is verifiable today with quoxproof, the offline verifier for tool-call receipts and a live Python package on PyPI. The hash-chain, timestamp and witness checks are what the VOLT and WARD protocols are built to support offline; a single command that walks all four layers of a full pack in one pass is still maturing, so verify the exact flow on your own instance before you rely on it in an audit.

  1. Install the verifier on a clean machine. pip install quoxproof. No account, no API key, no network dependency after install. This matters: an auditor who has to log into your console to verify your evidence has not verified anything independent.

  2. Check each receipt signature. Every tool-call receipt carries an Ed25519 signature over its contents. quoxproof confirms each signature against the public key in the pack, offline, today. A failed signature means the record was altered after signing. A missing signature means it is not evidence, it is a note.

  3. Walk the hash chain. Each VOLT ledger entry includes the hash of the previous entry. An offline check recomputes the chain end to end. If any record was inserted, removed, or reordered, the recomputed hash diverges and the check reports exactly where. This is the property a mutable log export cannot survive.

  4. Confirm the timestamps against an external anchor. RFC 3161 timestamps are issued by an independent time authority, so the auditor is not trusting your server clock. The check confirms that each record was stamped when it claims, in order.

  5. Confirm the witness receipts. The WARD receipts let the auditor confirm records existed at their stated time without exposing the underlying payloads, which is what lets you share proof with a customer's auditor without handing over customer data.

At the end the auditor has a pass or fail per record, produced on their hardware, that stands on cryptography rather than on your reputation. That is the difference between showing evidence and asserting compliance.

Mapping the pack to SOC 2 and the EU AI Act

SOC 2. An auditor testing controls over a period wants evidence the control operated every time, not a point-in-time snapshot. A hash-chained ledger of every agent action, with policy decisions attached, reads as continuous evidence for the monitoring and logical-access criteria. It answers "show me that access control and change monitoring operated across the audit window" with a record per event rather than a sampled screenshot.

EU AI Act. For high-risk systems, Article 12 requires automatic recording of events (logging) over the system's lifetime, and Article 19 requires those logs be kept. A tamper-evident, timestamped ledger of agent actions maps directly onto that record-keeping duty, and being able to verify it offline is what makes the record trustworthy to a national authority rather than only to you. Our reading of how this maps to agent systems is set out on the EU AI Act page. This is a mapping, not legal advice or a certification claim.

What exists today, and what is being built

Be precise about maturity, because compliance is a high-trust area. The open protocols (AEE, AOCL, VOLT, WARD) are shipped, and quoxproof is live on PyPI and verifies tool-call receipt signatures offline today. Quox itself is self-hosted (Docker), with per-organisation isolation and human-in-the-loop approvals on sensitive actions. The full control plane is pre-launch, so treat one-button pack assembly across all four layers as maturing rather than a finished product, and verify the exact export flow against your own instance before you rely on it in an audit. What you can run yourself now is the signature check via quoxproof; the hash-chain, timestamp and witness checks are the model the VOLT and WARD protocols are built to support, and the end-to-end pack verifier is maturing toward that.

The design principle underneath all of it is on the security page: evidence by default, so proof is a property of the system rather than a report someone assembles after the fact.