Orpheus TTS
Canopy Labs has introduced Orpheus, a family of state-of-the-art speech large language models (LLMs) designed for human-level speech generation. These models are built on the Llama-3 architecture and are trained on over 100,000 hours of English speech data, enabling them to produce natural intonation, emotion, and rhythm that surpasses current state-of-the-art closed source models. Orpheus supports zero-shot voice cloning, allowing users to replicate voices without prior fine-tuning, and offers guided emotion and intonation control through simple tags. The models achieve low latency, with approximately 200ms streaming latency for real-time applications, reducible to around 100ms with input streaming. Canopy Labs has released both pre-trained and fine-tuned 3B-parameter models under the permissive Apache 2.0 license, with plans to release smaller models of 1B, 400M, and 150M parameters for use on resource-constrained devices.
Learn more
GPT-Live
GPT-Live is a new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice. It is built to make talking with AI feel much more like having a real conversation through a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT-Live can show it is paying attention with short acknowledgments like “mhmm” or “yeah,” engage in quick back-and-forth, or stay quiet when the user needs a moment to think. Instead of processing separate turns one after another, GPT-Live continuously processes input while generating output, allowing it to decide many times per second whether to speak, keep listening, pause, interrupt, or invoke a tool. For questions that require web search, deeper reasoning, or more complex work, GPT-Live can delegate to a frontier model behind the scenes and bring the result back into the conversation when it is ready, while still maintaining the flow of the voice interaction.
Learn more
EVI 3
Hume AI's EVI 3 is a third-generation speech-language model that streams in user speech and forms natural, expressive speech and language responses. At conversational latency, it produces the same quality of speech as our text-to-speech model, Octave. Simultaneously, it responds with the same intelligence as the most advanced LLMs of similar latency. It also communicates with reasoning models and web search systems as it speaks, “thinking fast and slow” to match the intelligence of any frontier AI system. EVI 3 can instantly generate new voices and personalities instead of being limited to a handful of speakers. For instance, users can speak to any of the more than 100,000 custom voices already created on our text-to-speech platform, each with an inferred personality. No matter the voice, it responds with a wide range of emotions or styles, implicitly or on command.
Learn more
GPT-Live-1
GPT-Live-1 is one of the two new GPT-Live voice models rolling out to ChatGPT users globally, built to make talking with AI feel much more like having a real conversation. It is powered by a full-duplex architecture, so it can listen and speak at the same time instead of waiting for one rigid turn to end before the next begins. During conversations, GPT-Live-1 can show it is paying attention with short acknowledgments, engage in quick back-and-forth, pause when the user needs a moment to think, or stay quiet when asked to listen. It continuously processes input while generating output, allowing the model to decide many times per second whether to speak, keep listening, pause, interrupt, or invoke a tool. GPT-Live-1 also separates natural interaction from deeper work: when a question requires web search, reasoning, or more agentic capabilities, it can delegate the task to a frontier model behind the scenes and bring the result back when it is ready.
Learn more