Self hosted AI: what you actually control

Reason
AgentsQuoxMindAgentic TeamsMirrorQuoxLensRemember
QuoxMemoryBrain2CompoundingQuoxPlanCodebase MirrorAct
QuoxFlowQuoxEngineQuoxAgentQuoxChatAutonomyRun
EnterpriseOrganisationsPowers & ToolsmithQlusterQuoxBastionInterfaces
QuoxMCPQuoxCLIQuoxTerminalQuox ConsoleQuoxBoxGovern
HITL ApprovalsQuox SecurityAgent HonestyQuoxVaultAI GovernanceAgentic AIProve
For AuditorsVerifiable AI OpsLoggingQuoxSEOEU AI ActCompliance SuiteChannels
Matrix RoomsDiscord ProTelegram ProQuoxSignalCoreCommsAll products A-ZBuild
QuoxProofDeveloper KitPlugin SDKBring your tool to QuoxQuoxpertQuoxSkillsQlarityShip and sell
Build and sellBrowse MarketplaceDownloadsProtocolsDev Suite
Dev ServersQuoxBuildQuoxSlotsShared Skills + RulesDev workflow
RepoBrainGripeTriageFixLoopProofLoopDoneEngineAll products A-ZGet started
OverviewArchitectureProtocols
AEEAOCLVOLTWARDReference
GlossaryAPI ReferencePlugin SDKDockerAll products A-Z
It is 02:13. An agent investigating a failing service proposes a production restart. The useful question is who gave it permission, and what happens if nobody answers.
Imagine that incident on infrastructure you operate. The agent can read diagnostic material and propose a next step. You want help narrowing the fault, but a restart could interrupt work. Keeping the logs on your server matters. So does the mechanism that decides whether the agent may act.
Self-hosted AI often starts with a promise about data staying with you. That is worth having when the whole processing path stays local. For agents, though, the sharper distinction is what governs their actions. A process running in your rack can still have excessive permissions.
Here is the boundary plainly: self-hosting QuoxCORE's control plane does not mean everything runs locally. The language model can still be a cloud API, such as OpenAI or Anthropic. If you call a cloud model, the prompt goes to that provider. Local inference requires pointing Quox at a local or open-weights model you run yourself.
What stays on your infrastructure is the governance: policy, human approvals, spend budgets, memory and evidence. The decision about which model to call is yours too.
Follow the proposed restart through that distinction. This is a worked example of a control decision, not a report from a customer deployment.
The model has suggested an action. That suggestion needs to encounter a permission check before anything executes. In the agent loop, this is the consequential boundary: reasoning produces a proposal; permission determines whether execution may follow.
Quox evaluates policy before execution. Its human approvals have real deadlines, with no automatic approval on timeout. In our example, if a proposed action requires a human decision and the operator does not respond, silence cannot become permission.
The same distinction applies to cost. An agent can keep investigating without getting closer to an answer. Quox provides spend budgets at four scopes with a hard stop. The operator still has to choose sensible limits, but the limit has an enforcement mechanism. A request in a prompt to spend carefully is no substitute for that.
QuoxCORE is the self-hosted AI platform that runs these controls. It ships as a Docker stack containing authentication, a collector, a workflow engine, memory and agents, with a dashboard and a quox CLI. The policy engine and the records it produces run on your boxes.
That arrangement gives self-hosted AI agents a control plane you operate. You decide where to run it, maintain the infrastructure beneath it and retain its governance records. The policy, approval and budget controls are part of that deployment.
Credentials sit in a vault, encrypted with per-organisation keys. They are scoped by organisation so one team cannot read another team's credentials. That boundary matters when agents need access to tools, although you still need to decide how much access each credential should confer.
Return to the incident at 09:00. Someone asks whether the agent received permission to restart the service. A transcript showing that it wanted to restart is insufficient. You need the gate's decision.
Quox records every gate decision in an append-only evidence layer. AEE envelopes package those records; WARD witnesses them to make later alteration detectable. These are mechanisms for keeping tamper-evident evidence of the controls' decisions.
The distinction needs care. Tamper-evident does not mean tamper-proof. A signature is not truth: it cannot make an incorrect observation accurate. Nor does a recorded permission prove that the proposed action subsequently succeeded. You still need operational evidence from the affected service to establish what happened there.
The narrower benefit is useful enough. You retain a record of the gate decision on infrastructure you control, which gives the review something firmer than a reconstruction from chat.
Putting the stack on your server does not make the model correct. An agent can misread a log, trust misleading input or recommend a harmful change. Policy can restrict what it may do; it cannot turn every permitted action into a good decision. Human approval also depends on the reviewer understanding the request.
It does not make the deployment airgapped. If diagnostic text enters a prompt sent to a cloud API, that text leaves your infrastructure. Keeping memory locally does not change this. You need to inspect what you send, and choose local inference when your requirements demand it.
You also take on the ordinary work of running the stack: updates, backups, recovery and protecting the host and its keys. A self-hosted control plane gives you responsibility alongside control. It does not remove that work from the bill.
For the operator in this example, the gain is specific: a proposed production action meets controls they operate, an unanswered approval stays unanswered, and the gate's decision remains available for review. Model choice sits alongside those controls.
That is the practical case for self-hosted AI governance. QuoxCORE is free to run on your own box for a small footprint; paid tiers exist for larger fleets and enterprise plugins.
Start with the action you need to govern. Decide what permission it requires, who can grant it and how much the investigation may spend. Then choose where the model runs. Those are separate decisions, and self-hosting becomes more useful when you keep them explicit.