A small HTTP server that wraps Piper text to speech and streams the audio out at the rate a phone call actually plays it.
That sounds like a detail and it is the whole point. Synthesis produces audio far faster than real time. If you hand all of it to a telephony channel at once, the far end keeps a fraction of a second and discards the rest, and the caller hears the tail of a sentence with nothing before it. This server paces output instead, so what arrives is what gets heard.
It also handles the awkward parts of running Piper under load: the voice model is guarded so concurrent requests cannot corrupt each other, and a turn can be cancelled mid-sentence when a caller interrupts, rather than rendering audio nobody will hear.
Runs locally, so call audio does not leave your infrastructure. MIT licensed, on PyPI and Docker Hub.
Features
- Streams synthesis at the pace a call plays it, not as fast as it renders