Local voice AI: inference on your GPUs, data in your environment.
Local mode runs a full open-model stack of leading open models on GPU infrastructure you control. It is one half of the PeachDesk model plane, and it is the foundation of self-hosted voice AI on PeachDesk.
Local mode carries no per-minute provider charge for inference, because the models run on your GPUs instead of a metered third-party API. It is not free: you pay for GPU capacity, whether that is hardware you own or GPU infrastructure you rent. Frontier mode, the managed half of the platform, is metered per minute. Owning the capacity means your variable inference cost stops scaling with call volume, and every metered call in PeachDesk records its sourcing mode so the difference is visible in your own numbers. For the full breakdown, read voice AI cost per minute.
Sourcing is a per-stage decision, not a platform decision
A voice agent is ears, a brain, and a voice: speech-to-text, a language model, and text-to-speech. PeachDesk lets each stage choose its own sourcing per agent. Flip the presets to see the same pipeline three ways.
Every stage runs on PeachDesk-managed providers. Convenient, fully hosted, and metered per minute.
Each stage picks its own sourcing. Managed where you want convenience, your GPUs where volume justifies capacity. Sourcing can also include your own provider keys.
Every stage runs on leading open models on your GPUs. No per-minute provider charge; you pay for GPU capacity. Inference data never leaves your environment.
Built for operators
Operators control which models teams may use, load model weights offline so inference keeps running without outbound calls, and can scale GPU capacity to zero when the fleet is idle. Recordings and transcripts live in object storage the deployment controls, behind signed, time-limited URLs, and product telemetry is off unless you turn it on.
Design the agent these stages run in on the agent canvas, then choose the deployment options that match how much of the stack you want to own.
Own the stack. Do not just rent the pipeline.
Local mode is how the three PeachDesk pillars become architecture: own your cost, because inference spend becomes GPU capacity you control; own your data, because speech and language processing stay inside your environment; own your margin, because the cost floor of every call is yours to set. Two model planes, one platform.
Local mode runs a full open-model stack of leading open models on GPU infrastructure you control, covering the speech-to-text, language model, and text-to-speech stages. PeachDesk does not publish a fixed model list, because operators choose and approve which models their teams may use, and the stack is updated as open models improve.
Is Local mode really zero cost per minute?
Local mode carries no per-minute provider charge for inference, because the models run on your GPUs instead of a metered third-party API. It is not free: you pay for GPU capacity, whether that is your own hardware or rented GPU infrastructure. Frontier mode, by contrast, is metered per minute through managed providers.
Does Local mode work without an internet connection?
Yes, once the model weights are cached. Local mode supports offline model loading, so inference keeps running inside your environment with no outbound calls to model providers. Voice calls still depend on your telephony carrier connection, but the speech and language stages run entirely on your own GPUs.
Can I mix Local mode and Frontier mode in one agent?
Yes. Sourcing is chosen per pipeline stage and per agent. You might run speech-to-text and the language model on PeachDesk-managed providers while text-to-speech runs locally on your GPUs, or the reverse. This per-stage mix and match lets teams put local capacity exactly where call volume justifies it.
We use essential cookies to make PeachDesk work. Until you choose, analytics runs in a limited cookieless mode; accepting adds analytics cookies for fuller measurement. Read our cookie policy.