Advanced Voice plugin · content and media
Say it out loud. Hear it answer.
Full voice I/O for QuoxCORE. Hold to talk and OpenAI Whisper transcribes you in real time, the words become an intent, the intent runs through a governed action, and your agents answer through more than 40 ElevenLabs voices, each one assignable, tunable and yours to configure.
In plain words
What it is, where it lives, when to reach for it
- What is it
- A plugin for the QuoxCORE dashboard that gives agents better speech, with more than 40 professional voices to choose from.
- Where do I use it
- In chat, where agent replies can be read aloud; voices are picked under Settings, Voice.
- When would I use it
- When you want to hear agent replies read aloud in a voice you actually like.
- How do I use it
- Add it from the store, open Settings, choose Voice, pick a voice, and play any reply.
QuoxCORE is the free, self-hosted platform underneath this. What is QuoxCORE
Per-agent voice assignment · pick from 40+ ElevenLabs voices per agent
Press run: the question becomes an intent, the intent goes through a governed action, and the reply is spoken aloud in an ElevenLabs voice with the detail behind it printed underneath. In the real plugin your speech reaches the same pipeline through Whisper.
Voice output
Every agent gets its own voice.
The field behind this page is that same pipeline: it moves with the walkthrough below, or with your own voice if you opt in.
Assign a different voice to each of your agents, so you always know who is speaking. CommanderQ, WARDEN, CIPHER and every specialist can sound distinct, drawn from a library of more than 40 professional voices.
Fine control
Tune it until it sounds right.
Four parameters shape every voice: stability, similarity boost, style and speed, from 1.0x up to 1.5x. Drag the sliders and the waveform responds, the same way the voice would.
Three verbosity modes control how much gets spoken. Quiet keeps confirmations to 150 characters, Normal allows 500, Verbose 1,000. Premium raises the ceiling to 2,000 characters and adds AI-powered smart summarisation, so long answers are condensed for natural listening.
- Stability, similarity boost, style and speed controls
- Speed range 1.0x to 1.5x
- Quiet 150 · Normal 500 · Verbose 1,000 characters
- Premium: 2,000 character limit and smart summarisation
- Code blocks, tables, IPs and URLs stripped before speech
- Export your voice configuration and import it on another instance
Voice parameters · live
Under the hood
One pipeline, both directions.
SPEECH-TO-TEXT (STT) Input: Push-to-talk (hold mic button) Capture: MediaRecorder → WebM blob Endpoint: POST /webhook/transcribe Engine: OpenAI Whisper (whisper-1) Output: Transcribed text → chat input TEXT-TO-SPEECH (TTS) Voices: 40+ ElevenLabs · per-agent assignment Models: Flash · Turbo · Multilingual v2 Params: stability · similarity · style · speed 1.0-1.5x Limit: 2,000 chars (premium) · 500 (normal) Preprocess: strip code, tables, IPs, URLs ERRORS HANDLED HTTPS required · mic permissions · device availability missing API keys → clear messages point to Settings
Advanced Voice
no licence key required