Get started
Free Plugin

Local Inference
Your Hardware, Your Models, Your Data

Connect QuoxCORE to self-hosted LLM inference servers. Supports Ollama, vLLM, llama.cpp, and NVIDIA NIM. Zero API costs, full data privacy, complete audit trails.

Read the DocsView Store

Add it and an admin installs it from the dashboard; no checkout.

In plain words

What it is, where it lives, when to reach for it

What is it
A plugin for the QuoxCORE dashboard that adds Register self-hosted LLM servers (Ollama, vLLM, llama.cpp.
Where do I use it
In the QuoxCORE dashboard you run yourself, once the plugin is installed for your organisation.
When would I use it
When you want this capability available in your dashboard instead of in a separate tool.
How do I use it
Add it free from the Quox store, then an org admin installs it for your organisation.

QuoxCORE is the free, self-hosted platform underneath this. What is QuoxCORE

Local Inference. Your hardware, your models, your data. Zero egress, full audit trail.

ZERO COST

Your Models, Zero API Fees

Run AI inference on hardware you already own. Connect QuoxCORE to Ollama, vLLM, llama.cpp, or NVIDIA NIM servers and route chat, tool use, and embedding requests through your local network. No per-token costs, no rate limits, no vendor lock-in.

Data never leaves your network. Every request stays within your infrastructure boundary, making local inference the default choice for regulated industries, air-gapped environments, and teams that take data residency seriously.

  • Ollama, vLLM, llama.cpp, NVIDIA NIM
  • Auto-detect backend type
  • OpenAI-compatible API
  • SSE streaming support
CAPABILITIES

Built for Local Inference

Server Connection

Auto-detect and connect to inference servers on your network. Supports multiple backend types through a single configuration interface.

Model Switching

Switch between locally loaded models on the fly. See available models, VRAM usage, and context window sizes before selecting.

Streaming

Real-time SSE streaming responses from local models. Same streaming experience as cloud providers, with none of the latency penalties.

Embeddings

Local embedding generation for RAG pipelines. Keep your vector data on-premises alongside the models that produce it.

Health Monitoring

Automatic health checks and status reporting for connected inference servers. Detect failures before they affect users.

Usage Metrics

Token usage, latency percentiles, and request tracking per model and per server. Understand your inference workload without external tooling.

PRICING

Choose Your Tier

All tiers are free during the public beta. Launch pricing shown for reference.

Free
For hobbyists and solo developers
You run Ollama on a workstation or spare machine. You want to try local inference without committing to anything.
$0Free
Free forever
1 inference source
Auto-detect backend type
Basic streaming responses
Manual health check
Model switching
Embedding support
Usage metrics
AEE/VOLT audit trail
GPU fleet management
MOST POPULAR
Pro
For teams with dedicated GPU servers
Your team runs vLLM or llama.cpp on one or more machines. You need multiple models, embeddings, and visibility into what your inference stack is doing.
$89Free
Free during beta · $89 one-time at GA
Unlimited inference sources
Auto-detect backend type
Full SSE streaming
Auto health checks (30s)
Model switching on the fly
Local embedding generation
Token and latency metrics
Data residency labels
GPU fleet management
Enterprise
For organisations running GPU fleets
You operate multiple GPU machines across racks, sites, or data centres. You need fleet-wide visibility, routing, RBAC, and audit evidence that satisfies your compliance team.
$699Free
Free during beta · $699 one-time at GA
Everything in Pro
QuoxAgent GPU plugin per host
Auto-discovery of GPU nodes
VRAM and temperature monitoring
Multi-node inference routing
Model allowlists per team
Local-to-cloud failover policies
AEE/VOLT audit per inference
Enforced data residency
All tiers are free and fully unlocked during the public beta. Tier limits activate at general availability. Early adopters who provide feedback during beta will receive a discount on launch pricing.
ENTERPRISE

GPU Fleet Management

Install QuoxAgent on GPU machines to enable automatic discovery, VRAM monitoring, and temperature tracking across your inference fleet. QuoxCORE sees every GPU node, its loaded models, and its current utilisation in real time.

Multi-node inference routing distributes requests across available servers based on model availability, queue depth, and hardware health. Define model allowlists per team, enforce RBAC policies, and maintain full AEE/VOLT audit trails for every inference request.

  • QuoxAgent GPU plugin
  • VRAM and temperature monitoring
  • Multi-node inference routing
  • Model allowlists and RBAC
  • AEE/VOLT audit per inference
  • Data residency enforcement
AUDIT

Every Inference, Verified

AEE Envelopes

Every inference request produces a signed AEE envelope capturing the model, prompt hash, token counts, and latency. Compliance without manual logging.

VOLT Recording

Tamper-evident inference records written to your VOLT ledger. Prove what model produced what output, and when.

Privacy Boundary

Data residency tags per inference source. Enforce that sensitive workloads never route to external endpoints, even as a fallback.

Failover Policy

Configurable local-to-cloud fallback rules. Define what happens when local servers are unavailable: queue, reject, or route to an approved cloud endpoint.

Run AI on your own terms

Local Inference connects QuoxCORE to the models running on your hardware. No API keys, no per-token billing, no data leaving your network. Every inference request is recorded in an AEE envelope for full auditability.

Full Documentation

Stay updated

Product updates, new features, and protocol news.