Local Inference
Your Hardware, Your Models, Your Data
Connect QuoxCORE to self-hosted LLM inference servers. Supports Ollama, vLLM, llama.cpp, and NVIDIA NIM. Zero API costs, full data privacy, complete audit trails.
Add it and an admin installs it from the dashboard; no checkout.
In plain words
What it is, where it lives, when to reach for it
- What is it
- A plugin for the QuoxCORE dashboard that adds Register self-hosted LLM servers (Ollama, vLLM, llama.cpp.
- Where do I use it
- In the QuoxCORE dashboard you run yourself, once the plugin is installed for your organisation.
- When would I use it
- When you want this capability available in your dashboard instead of in a separate tool.
- How do I use it
- Add it free from the Quox store, then an org admin installs it for your organisation.
QuoxCORE is the free, self-hosted platform underneath this. What is QuoxCORE
Local Inference. Your hardware, your models, your data. Zero egress, full audit trail.
Your Models, Zero API Fees
Run AI inference on hardware you already own. Connect QuoxCORE to Ollama, vLLM, llama.cpp, or NVIDIA NIM servers and route chat, tool use, and embedding requests through your local network. No per-token costs, no rate limits, no vendor lock-in.
Data never leaves your network. Every request stays within your infrastructure boundary, making local inference the default choice for regulated industries, air-gapped environments, and teams that take data residency seriously.
- Ollama, vLLM, llama.cpp, NVIDIA NIM
- Auto-detect backend type
- OpenAI-compatible API
- SSE streaming support
Built for Local Inference
Server Connection
Auto-detect and connect to inference servers on your network. Supports multiple backend types through a single configuration interface.
Model Switching
Switch between locally loaded models on the fly. See available models, VRAM usage, and context window sizes before selecting.
Streaming
Real-time SSE streaming responses from local models. Same streaming experience as cloud providers, with none of the latency penalties.
Embeddings
Local embedding generation for RAG pipelines. Keep your vector data on-premises alongside the models that produce it.
Health Monitoring
Automatic health checks and status reporting for connected inference servers. Detect failures before they affect users.
Usage Metrics
Token usage, latency percentiles, and request tracking per model and per server. Understand your inference workload without external tooling.
Choose Your Tier
All tiers are free during the public beta. Launch pricing shown for reference.
GPU Fleet Management
Install QuoxAgent on GPU machines to enable automatic discovery, VRAM monitoring, and temperature tracking across your inference fleet. QuoxCORE sees every GPU node, its loaded models, and its current utilisation in real time.
Multi-node inference routing distributes requests across available servers based on model availability, queue depth, and hardware health. Define model allowlists per team, enforce RBAC policies, and maintain full AEE/VOLT audit trails for every inference request.
- QuoxAgent GPU plugin
- VRAM and temperature monitoring
- Multi-node inference routing
- Model allowlists and RBAC
- AEE/VOLT audit per inference
- Data residency enforcement
Every Inference, Verified
AEE Envelopes
Every inference request produces a signed AEE envelope capturing the model, prompt hash, token counts, and latency. Compliance without manual logging.
VOLT Recording
Tamper-evident inference records written to your VOLT ledger. Prove what model produced what output, and when.
Privacy Boundary
Data residency tags per inference source. Enforce that sensitive workloads never route to external endpoints, even as a fallback.
Failover Policy
Configurable local-to-cloud fallback rules. Define what happens when local servers are unavailable: queue, reject, or route to an approved cloud endpoint.
Run AI on your own terms
Local Inference connects QuoxCORE to the models running on your hardware. No API keys, no per-token billing, no data leaving your network. Every inference request is recorded in an AEE envelope for full auditability.
Stay updated
Product updates, new features, and protocol news.