Skip to main content
An optional, self-contained subsystem for running speech and language models on your own hardware. One gateway on :8100 fronts three slots — STT, TTS, LLM — and each slot runs whichever model you selected, without the gateway knowing which one.
The model server is optional. VoicEra runs fine against cloud providers alone. Deploy this when data residency, per-minute cost, or Indic-language coverage makes self-hosting worth the GPUs.

Start here

The models

Operating it

Honest status

Two gaps are documented rather than hidden. The LLM slot has not yet been run on real hardware, and a full call routed from the voice runtime through a self-hosted model has not been verified end to end. Both are called out on the pages concerned.