Alternatives to mrmr
Compare mrmr alternatives for your business or organization using the curated list below. SourceForge ranks the best alternatives to mrmr in 2026. Compare features, ratings, user reviews, pricing, and more from mrmr competitors and alternatives in order to make an informed decision for your business.
-
1
Telnyx
Telnyx
Telnyx is a global communications infrastructure platform that provides voice, messaging, networking, and AI-powered real-time communication capabilities through a fully owned telecom stack. The platform combines carrier-grade networking, programmable identity systems, AI inference, and low-latency communication infrastructure to support real-time conversational AI agents and enterprise communication workflows. Telnyx owns and operates its entire network stack, including physical infrastructure, mobile core systems, edge processing, and AI compute layers, enabling faster performance and lower latency without relying on third-party telecom providers. The platform offers tools such as voice agent builders, speech-to-text, text-to-speech, global phone numbers, AI orchestration, and programmable compliance controls for building intelligent voice and messaging systems. -
2
CSS IMPACT
CSS, Inc
CSS IMPACT is a leading provider of Next-Gen Financial Ecosystems & Omnichannel Engagement cloud platforms. Featuring HD 2.0 | Ai - an Agent-less “Ai” (Artificial Intelligence) Digital Consumer or Debtor Engagement bot for credit, billing, collections & revenue cycle management. This “Digital & Voice First Ai" servicing technology can answer common questions, accept payments, & negotiate accounts with a frictionless positive experience without changing the consumer's behavior by using new IoT channels of communications, such as Google Assistant, Google Ai Voice (phone), Text, Chat, Email, Online smart portals, as well as traditional call center technologies - Dialers, Click-to-dial, IVR & Telephones. -
3
Dialogflow
Google
Dialogflow from Google Cloud is a natural language understanding platform that makes it easy to design and integrate a conversational user interface into your mobile app, web application, device, bot, interactive voice response system, and so on. Using Dialogflow, you can provide new and engaging ways for users to interact with your product. Dialogflow can analyze multiple types of input from your customers, including text or audio inputs (like from a phone or voice recording). It can also respond to your customers in a couple of ways, either through text or with synthetic speech. Dialogflow CX and ES provide virtual agent services for chatbots and contact centers. If you have a contact center that employs human agents, you can use Agent Assist to help your human agents. Agent Assist provides real-time suggestions for human agents while they are in conversations with end-user customers. -
4
11.ai
ElevenLabs
11.ai is a voice-first AI assistant built on ElevenLabs Conversational AI that connects your voice to everyday workflows via the Model Context Protocol (MCP), enabling hands-free planning, research, project management, and team communication. By integrating out of the box with tools such as Perplexity for live web research, Linear for issue tracking, Slack for messaging, and Notion for knowledge management, and supporting custom MCP servers, 11.ai can interpret sequential voice commands, contextualize data, and take meaningful actions. It delivers real-time, low-latency interactions with multimodal support (voice and text), integrated retrieval-augmented generation, automatic language detection for seamless multilingual conversations, and enterprise-grade security (including HIPAA compliance). -
5
AnveVoice
AnveVoice
AnveVoice is an AI-powered voice agent platform that turns websites into interactive, conversational experiences. It allows businesses to deploy intelligent voice assistants that can talk to visitors, answer questions, guide navigation, and complete actions like form filling and lead capture in real time. Unlike traditional chatbots, AnveVoice is voice-first and action-driven—helping businesses increase conversions, reduce drop-offs, and automate customer interactions without manual support. With plug-and-play integration, multilingual capabilities, and no-code setup, companies can launch a fully functional AI voice assistant on their website in minutes.Starting Price: $39/month -
6
OpenWorker
OpenWorker
OpenWorker is an open source, local-first AI coworker that gets everyday tasks done from start to finish instead of only returning answers. Users ask for an outcome, such as a renewal brief, incident report, follow-up message, calendar update, sprint summary, or finished document—and OpenWorker works across the tools where the information already lives. It can connect with Slack, Gmail, Outlook, Google Calendar, Notion, HubSpot, GitHub, Attio, Google Drive, Jira, Linear, Asana, Dropbox, Box, files, and other services through one-click or manual connections. It supports cloud, open-weight, and fully local models, including providers such as OpenAI, Anthropic, Google, xAI, Mistral, DeepSeek, Kimi, Qwen, and Ollama, and users can switch models when a task calls for something different. OpenWorker researches, gathers context, performs multi-step work, creates polished outputs in chat, Slack, Markdown, PDF, images, or files, and checks in before consequential actions.Starting Price: Free -
7
Floatbot
Floatbot.AI
Floatbot.AI is a powerful Voice-First, Multi-Modal Conversational AI + Co-Pilot Platform Floatbot.AI is a Multi-Modal Conversational AI (Voice first) + Co-Pilot Platform designed to supercharge operations in Insurance, Collections, Lending, Banking, and BPOs. From redefining customer engagement, streamlining processes to empowering agents and employees, we are your partner in driving smarter, faster and impactful business interactions. With our no-code/low-code platform, you can build powerful AI Agents in minutes—no technical expertise required. Floatbot.AI is trusted by 200+ top players in insurance, banking, & collections to innovate and scale customer engagement & operational excellence.Starting Price: $99 -
8
Mumble AI
Mumble AI
Mumble AI is a voice-first productivity app that replaces the need for separate meeting recorders, note-taking apps, and dictation tools. Everything runs in one Mac app with both local and cloud AI built in. Mumble connects the full voice workflow in one place. You can record a meeting, capture a quick idea by voice, or dictate a full email, all without switching tools. Local mode keeps everything on your Mac for privacy. Cloud mode delivers higher accuracy across 40+ languages. You can switch anytime. Key Features No-Bot Meeting Recording Captures audio directly from your Mac with no bot joining the call. Works with Zoom, Google Meet, Teams, Slack, and any app that plays audio through your Mac. Live transcription with speaker labels so you can see who said what in real time. Instant structured summaries the moment your meeting ends. Auto-detects meetings from your Google Calendar and can automatically start and end recording for you. -
9
Cal.ai
Cal.ai
Cal.ai adds AI-powered voice agents to the Cal.com scheduling platform so that phone calls, reminders, confirmations, follow-ups, booking calls, and no-shows can be automated using natural, human-like agents. You can set up triggers based on events in your existing workflows (for example, on form submissions, meeting no-shows, cancellations), assign a phone number for the AI agent to call from (you can import an existing one), and write custom prompts to control the tone, personality, and script of each interaction in voice. The system also integrates deeply with Cal.com’s calendar syncing (Google, Outlook, etc.), scheduling links, team scheduling, group meetings, and route bookers to the right person based on availability and event type. Calls include analytics; transcripts, completion rates, booking outcomes, sentiment/tone detection, and other performance metrics to help you refine conversations and improve conversion.Starting Price: $0.29 per minute -
10
ECHO by Zencia AI
Zencia AI
ECHO by Zencia is a SaaS platform for building, deploying, and managing production-ready AI voice agents. Create AI receptionists, sales agents, customer support assistants, recruiters, or custom AI voice employees without the complexity of integrating telephony, speech-to-text, large language models, text-to-speech, and workflow automation from scratch. ECHO combines persistent memory, custom knowledge bases, knowledge-gap detection, and intelligent workflows to deliver natural, context-aware voice conversations. Connect your CRM, calendars, and business tools to automate inbound and outbound calls, qualify leads, schedule appointments, answer customer queries, and execute business actions from a single dashboard. With multilingual support, analytics, call history, and centralized agent management, ECHO enables startups, SMBs, and enterprises to deploy scalable Voice AI that remembers context, takes action, and helps automate business communication. -
11
Leadlock
Leadlock
Leadlock is a speech-to-speech voice AI platform built specifically for GoHighLevel agencies, helping them answer every call, qualify leads, book appointments, and update GHL pipelines in real time. Unlike traditional voice AI stacks that chain speech-to-text, an LLM, and text-to-speech, it supports true multimodal speech-to-speech through OpenAI Realtime and Gemini Live, alongside xAI Grok and ElevenLabs options, enabling sub-second latency, natural turn-taking, and interruptions. Agencies can choose from more than 72 voices across multiple providers and select different models for different agents and use cases. Native GoHighLevel integration connects contacts, calendars, pipelines, opportunities, tags, custom fields, workflows, and sub-accounts without middleware or Zapier-style glue. Before answering, agents can pull caller history and CRM context to personalize conversations from the first ring.Starting Price: $97 per month -
12
Gemini 3.8 Live
Google
Gemini 3.8 Live is Google DeepMind’s real-time speech-to-speech AI model for building conversational voice applications and interactive agents. The model can maintain natural dialogue while reasoning, using tools, and carrying out tasks during an ongoing conversation. Asynchronous function calling allows applications to execute API and tool requests in the background while Gemini continues streaming audio responses to the user. Gemini 3.8 Live can also incorporate live visual context, enabling agents to respond based on what users say and what the system can see. It supports more than 97 languages, maintains accent consistency, and is designed to accurately interpret alphanumeric information such as confirmation codes, claim numbers, and technical data. Gemini 3.8 Live is available through the Gemini Live API and Google AI Studio for developers building customer service agents, assistants, training applications, and other voice-first experiences. -
13
April
April
April is a voice-powered AI executive assistant that enables hands-free management of email and calendars, whether you're commuting, walking, or working out, allowing you to achieve Inbox Zero using natural voice commands. It intelligently summarizes long email threads, lets users dictate and send replies on the go, fetches meeting locations or Google Meet links from your calendar or inbox when you need them, and swiftly deletes thousands of promotional emails to declutter your inbox. Designed with secure, bank‑grade encryption and adaptive learning, it understands executive communication styles, grasps context and urgency, and continuously refines its understanding of your tone and preferences. Optimized for seamless use via AirPods, CarPlay, and Face ID, April transforms routine email and calendar workflows into effortless, voice-first interactions, helping busy professionals stay productive and organized without needing hands or screens.Starting Price: $29 per month -
14
Dograh
Dograh
Dograh is an open source, self-hostable voice agent platform with a no-code workflow builder for creating production voice agents. Teams can choose their own inbound channels, speech-to-text, LLM, text-to-speech, and telephony providers, or replace the traditional cascade with speech-to-speech models for direct audio-in, audio-out conversations with natural turn-taking, interruption handling, and low latency. The platform supports inbound and outbound calling, widgets, telephony integrations, observability, traces, real-time analytics, hybrid pre-recorded voice plus TTS, and more than 70 languages. Its MCP server lets Claude Code, Cursor, OpenClaw, Codex, and other agent runtimes create, modify, and deploy voice agents directly from development environments. Dograh can run on your own servers, inside a private cloud or VPC, or in a managed environment, with support for models hosted entirely within your perimeter.Starting Price: 1¢ per minute -
15
Tunk.ai
Tunk.ai
Tunk.ai is a multilingual AI voice agent platform that enables businesses to build, deploy, and manage intelligent voice agents for customer interactions and business workflows. With its Anchor agent builder, businesses can create AI voice agents using prompts and connect them to telephony, APIs, CRMs, MCP servers, external tools, and business systems. Tunk.ai supports speech-to-text, text-to-speech, real-time voice conversations, SIP and telephony integrations, tool calling, function calling, and multiple AI and speech providers. Agents can use connected tools to retrieve information, perform actions, update systems, and automate workflows. Tunk.ai can automate customer support, sales, lead qualification, appointment scheduling, recruitment, healthcare workflows, collections, reminders, and contact-center operations. The platform supports multiple languages, customizable workflows, APIs, integrations, and scalable infrastructure for SMB and enterprise deployments.Starting Price: $0.05 per minute -
16
Google has released updated Gemini audio models that significantly expand the platform’s capabilities for natural, expressive voice interactions and real-time conversational AI with the introduction of Gemini 2.5 Flash Native Audio and improved text-to-speech technology. The updated native audio model powers live voice agents that can handle complex workflows, follow detailed user instructions more reliably, and maintain smoother multi-turn conversations by better recalling context from previous turns. It is now available across Google AI Studio,Gemini Enterprise Agent Platform, Gemini Live, and Search Live, enabling developers and products to build interactive voice experiences such as intelligent assistants and enterprise voice agents. In addition to the real-time voice improvements, Google enhanced the underlying Text-to-Speech (TTS) models in the Gemini 2.5 family to offer greater expressivity, tone control, pacing adjustments, and multilingual support.
-
17
Boson AI
Boson AI
Boson AI provides voice agents powered by foundation audio models, built to run in business workflows and learn from every call. Higgs Realtime enables live voice agents for support lines, sales calls, and product assistants that listen, reason, call tools, and respond in real time with low latency and natural speech-to-speech interaction. Higgs Audio and Avatar extend these capabilities with text-to-speech, speech-to-text, voice cloning, sentiment detection, and avatar generation, producing natural speech while understanding tone, emotion, and intent. The models support high-accuracy multilingual speech recognition, real-time translation, and expressive voice generation, while sentiment signals can improve routing, analytics, and context-aware agent behavior. Designed for real-world production, the platform emphasizes quality, latency, reliability, and flexible deployment across managed and self-serve environments. -
18
Fluent
Fluent For All, Inc.
Fluent is a voice-first control layer for Windows 10 and 11. You speak or type a sentence describing the outcome you want; Fluent reads the window through the same accessibility tree that screen readers use, works out the steps, and invokes the real controls. No screenshots, no guessing at pixel coordinates, and no per-application plugin, so it drives software it has never seen before. Speech recognition runs on the device. Destructive steps stop and ask before they run. Every action is written to an append-only, hash-chained audit log, so an administrator can reconstruct exactly what happened. Built for people who cannot comfortably use a mouse or keyboard - strain injury, tremor, low vision, limited hand mobility - and for regulated desktops where screen content must not leave the machine unchecked. Speech, dictation, text-to-speech, hotkeys and accessibility navigation keep working whatever the licence state. 30-day trial.Starting Price: $20/month/user -
19
Grok Voice Agent Builder
SpaceXAI
Grok Voice Agent Builder is xAI’s no-code platform for configuring production voice agents on Grok Voice in under two minutes. It is built for operators and developers who want high-volume voice agents without building the surrounding stack from scratch, bringing telephony, knowledge retrieval, tools, guardrails, MCPs, and observability into one place. Instead of stitching together separate speech-to-text, language model, and text-to-speech APIs, Voice Agent Builder uses one interface on a speech-to-speech path built for Grok Voice, tightly coupled to the model rather than assembled from three different systems. Users can write a plain-language description of how calls should flow, attach documents, connect tools, set guardrails, and move quickly from zero to a working agent. It can retrieve from uploaded knowledge bases in common formats such as plain text, Markdown, Word, PowerPoint, Excel, HTML, JSON, and others.Starting Price: $30 per month -
20
Gemini 3.1 Flash Live
Google
Gemini 3.1 Flash Live is Google’s most advanced real-time audio model, designed to deliver natural, reliable, and low-latency voice interactions for the next generation of conversational AI. It is optimized for real-time dialogue, enabling fluid, human-like conversations with improved precision, faster response times, and a more natural rhythm that better reflects how people actually speak. It enhances tonal understanding, allowing it to recognize nuances such as pitch, pace, and emotional cues, and dynamically adapt responses to user intent, including frustration or confusion. Built for both developers and enterprises, it can be accessed through the Gemini Live API in Google AI Studio, as well as integrated into production environments to power voice-first agents capable of handling complex, multi-step tasks at scale. It supports multimodal inputs including text, audio, images, and video, and produces both text and audio outputs, enabling richer, context-aware interactions. -
21
Calldock
Calldock
Calldock is an AI-powered voice agent platform that instantly calls your website visitors when they leave a number with no forms, no waiting. Built for SaaS, service businesses, real estate, and more, it helps convert leads by answering questions, qualifying prospects, booking meetings, and syncing with tools like Slack, Zapier, and Google Calendar. You can fully customize agent behavior, voice, and call logic no code required. Get transcripts, intent detection, analytics, and up to 10 agents per account. Get your website a voice agent that sits like a chatbot, live in minutes and turn passive visitors into high-intent conversations without hiring extra reps.Starting Price: $49/month -
22
Sarvam Indus
Sarvam
Indus is Sarvam’s official conversational AI interface designed to give users direct access to its flagship sovereign language models through a simple, real-time chat experience. Introduced in February 2026 as a limited beta product, it serves as the primary interface for interacting with Sarvam’s 105-billion-parameter model, bringing advanced reasoning, multilingual understanding, and conversational capabilities into a single application. It is built to deliver an AI experience tailored specifically to Indian users, supporting more than 22 Indian languages, including native scripts and code-mixed inputs, while maintaining contextual understanding aligned with local culture and communication patterns. It enables both text and voice interactions, allowing users to speak naturally and receive responses in text or synthesized speech, creating a voice-first, accessible interface for diverse use cases. -
23
VoiceBun
VoiceBun
VoiceBun is an open source, no-code voice-agent builder that lets you create, configure, and deploy AI-powered conversational assistants entirely via natural-language prompts. It combines speech-to-text, large-language models, and text-to-speech into a unified platform where you define your agent’s goals, initial greeting, tool integrations and data sources; VoiceBun automatically generates the underlying conversational logic, state management and API connectors needed to handle inbound and outbound calls for support, scheduling, lead qualification and more. The web-based interface gives you mobile-friendly access and isolated deployments through user-specific subdomains, while built-in analytics surface call transcripts, usage metrics, success rates, and sentiment trends. Integration includes options for telephony, webhook actions for external workflows, and role-based access controls with encrypted credentials for enterprise security.Starting Price: $20 per month -
24
ElevenAgents
ElevenLabs
ElevenLabs Agents is a platform for building, deploying, and scaling intelligent conversational AI agents that can speak, type, and take action across phone, web, and application environments. It enables developers and teams to create real-time agents that interact naturally with users through voice and text, combining speech-to-text, large language models, and text-to-speech into a unified system that functions like a human conversation partner. It allows agents to resolve customer issues, automate workflows, answer questions, and execute tasks based on connected data sources and predefined logic, making interactions both accurate and context-aware. These agents can be customized with knowledge bases, system prompts, and tools that enable them to access external systems, execute custom logic, and perform actions beyond simple responses. They support multimodal capabilities, meaning they can read, speak, and interpret inputs while handling conversational dynamics.Starting Price: $5 per month -
25
Sidekick
Sidekick
Sidekick enables users to build powerful, Zapier-style automations simply through a conversational interface, no complex UI navigation required. You begin by describing what you want in plain language, and Sidekick’s AI automatically creates the workflow, visualizes it on a canvas, handles error logic, and lets you run or schedule the automation immediately. It integrates seamlessly with a range of everyday applications, such as Gmail, Google Calendar, Google Docs, Google Sheets, Notion, Airtable, HubSpot, Slack, and Linear, offering pre-built templates that you can customize via chat to match your workflow needs. Use cases include syncing Gmail emails to Google Sheets, summarizing calendar events and sharing them via Slack, storing inbound leads from email into Notion databases, automatically generating post-meeting documents, crafting weekly pipeline risk reports from HubSpot deals, creating Linear issues from spreadsheet entries, and delivering prioritized email digests.Starting Price: $19 per month -
26
TrustClaw
Composio
TrustClaw is a 24/7 AI assistant with 1000+ integrations via OAuth and sandboxed execution, built on the ideas behind OpenClaw and rebuilt from scratch with security at the foundation. It is designed as an AI that does things while you sleep; users can chat with the same agent across messaging apps like Telegram, with WhatsApp, Discord, and Slack listed as coming soon, and ask it to handle real workflows across connected tools. TrustClaw can fetch and categorize emails, draft replies, log customer complaints in Notion, summarize Slack messages, pull completed Linear tickets and draft release notes, scrape reviews, analyze sentiment, check Gmail for customer emails, and work across apps such as Gmail, GitHub, Notion, Figma, Linear, Jira, Google Drive, Google Calendar, Todoist, Asana, Trello, Stripe, HubSpot, Airtable, and many more. Its main promise is replacing unsafe password- or API-key-based agent setups with OAuth-only connections, encrypted managed credentials, etc.Starting Price: Free -
27
Oravoice
Oravoice
Oravoice – 24/7 AI Voice Agents That Never Miss a Lead. Every missed call is a missed opportunity. Oravoice deploys human-like AI voice agents that answer every call instantly, qualify leads, and book appointments automatically — 24/7, without hiring extra staff. Trusted across 30M+ conversations and saving businesses 2.5M+ hours, Oravoice handles the phones so your team can focus on closing deals. Key Features: 📞 24/7 inbound call handling with IVR-style routing 📤 Outbound call automation with CSV/Excel contact upload 📅 Auto appointment booking via Calendly & Google Calendar 💬 SMS follow-up automation 🎙️ 100+ natural, human-sounding voices across languages 📊 Call transcripts, data capture & CRM sync Perfect for: Clinics, Real Estate, Hospitality, Restaurants, Automotive, E-Commerce, Sales Teams Results businesses see: 40% faster response times 99.95% call availability Zero missed leadsStarting Price: $99/month -
28
Voksha
Voksha
Voksha is a natural voice AI receptionist purpose-built for small businesses. It answers every incoming call with sub-200ms human-like voice responses, books appointments directly into your existing calendar, captures and qualifies leads with structured data, and automatically recovers missed calls within 90 seconds. Key capabilities include: - 24/7 call answering with natural voice conversations - Instant appointment booking and calendar integration - Lead capture and qualification - Missed call recovery in under 90 seconds - Intelligent call forwarding based on availability and expertise - HIPAA compliance and social engineering immunity - Integration with Salesforce, HubSpot, Clio, Google Calendar, Slack, and more Starting at $15/month, no contracts, and setup in under 5 minutes.Starting Price: $15/month -
29
Vocode
Vocode
Vocode is an open source library that simplifies the creation of voice-based applications leveraging large language models. Developers can build real-time streaming conversations with LLMs and deploy them to phone calls, Zoom meetings, and more. Vocode provides easy abstractions and integrations so that everything you need is in a single library. It offers out-of-the-box integrations with leading speech-to-text and text-to-speech providers, including AssemblyAI, Deepgram, Google Cloud, Microsoft Azure, and Whisper. The platform supports cross-platform deployment across telephony, web, and Zoom, enabling applications like LLM-powered phone calls, personal assistants, and voice-based games. Vocode's modular design allows for seamless integration of various AI models and services, providing developers with the flexibility to choose the best components for their applications. The platform also supports multilingual capabilities.Starting Price: Free -
30
GAIA
GAIA
GAIA is an open source, proactive personal AI assistant that manages work across the tools people already use. It watches inboxes, calendars, tasks, and connected apps, surfaces what matters, and can act before the user asks. GAIA connects with Gmail, Slack, Notion, Google Calendar, GitHub, HubSpot, Todoist, Linear, and services through native integrations and MCP. Users can ask it to triage email, draft replies, schedule or reschedule meetings, prepare briefings, create tasks from conversations, conduct research, generate documents, and run multi-step workflows. Automations are described in plain language, then GAIA builds the steps, assigns a schedule or event trigger, and repeats the workflow across the connected stack. Its task system breaks requests into steps, sets priorities, performs research or drafting, notifies teammates, and closes work when execution finishes.Starting Price: $30 per month -
31
OmniDimension
OmniDimension
OmniDimension is a Conversational AI platform that helps businesses automate customer conversations across voice, web, and WhatsApp. Users can create AI voice agents for lead generation, appointment booking, customer support, collections, outbound calling, and inbound call handling without coding. The platform includes built-in telephony, AI-powered outbound campaigns, CRM integrations, workflow automation, knowledge base support, multilingual conversations, and real-time analytics. Businesses can connect tools like HubSpot, Salesforce, Google Calendar, Slack, Zapier, Make, n8n, and custom APIs to automate customer interactions and business workflows. OmniDimension enables organizations to deploy AI agents in minutes, scale customer engagement 24/7, and improve response times while reducing operational costs. -
32
Guzli
Guzli
Guzli is an AI chat and voice bot. It answers people on your website, lets them talk on the page, picks up the phone, and can call people back. Same answers. No coding. Works with Shopify, Stripe, Calendly, Zendesk, HubSpot, Salesforce, Slack, WhatsApp.Starting Price: $39/month -
33
WarmLane
WarmLane
WarmLane is a 24/7 AI voice agent built for businesses that lose money every time the phone goes unanswered. Every missed call is a customer who called the next business instead. WarmLane answers on the first ring — at 2 AM, on a Sunday, or while the owner is under a sink — and handles the call the way a well-trained receptionist would. It answers questions about services, hours, pricing and service area using your own business information, qualifies the caller, captures their name, number and reason for calling, and books the appointment directly onto your Google Calendar or Calendly. The caller gets a confirmation by text and email, which also cuts no-shows. When a call genuinely needs a person, WarmLane forwards or warm-transfers it to you or your team, so nobody is ever left stranded. Missed and dropped calls get an automatic callback and text. Every conversation is transcribed, searchable, and exportable.Starting Price: $79/month -
34
Vision Agents
Stream
Vision Agents is an open source Python framework for building low-latency voice and video AI agents with any model. It lets developers plug in LLM, speech, and vision models from more than 25 providers and ship real-time agents for telehealth, voice support, live coaching, video analysis, interactive avatars, security monitoring, sports commentary, and other multimodal applications. It is designed to help teams build agents that can listen, speak, see, process media, call tools, and respond in real time while running on Stream’s global edge network with sub-500ms latency. Developers can build a first agent in minutes, using a small Python setup with Gemini Realtime, OpenAI, Deepgram, ElevenLabs, Stream, or other supported providers. Vision Agents supports both real-time speech-to-speech models and custom STT/LLM/TTS pipelines, giving teams either the fastest path to a working voice agent or full control over speech recognition, language reasoning, text-to-speech, etc.Starting Price: Free -
35
Claudia by Neuro
Neuro Notion
Neuro is the world’s first AI personal assistant for adults with ADHD. Think Jarvis from Iron Man, but for ADHD Brains. It’s a voice-first platform that turns chaotic thoughts (via voice conversation) of an ADHDer into streamlined organization. All the user does is braindump and the system handles the rest for them, showing them what they need to know when they need to know it. -
36
Sauna
Sauna
Sauna is the first multiplayer AI: a cloud-based workspace that learns how you work, remembers everything that matters, and acts on it for you and your whole team. It is built for the reality that work lives everywhere, scattered across Google Sheets, documents, Linear, GitHub, Slack, Notion, meeting notes, Gmail, Google Calendar, and more. Sauna connects it all and puts it to work, so every task, decision, and workflow can be handled automatically from one mission control. It runs in the cloud around the clock, drafting messages in your voice, tracking what matters, filing tickets, delivering briefings, and keeping work moving even when you are offline. Sauna is also designed for team knowledge, not just individual productivity: every Sauna can open the door to another, so users can ask a colleague’s Sauna a question, pull their perspective, or tap into their knowledge directly without interrupting them.Starting Price: $99 per month -
37
Krybe
Krybe
Krybe is an AI-powered platform offering cutting-edge voice and transcription solutions, including voice agents and speech AI, designed to transform noise into actionable insights for businesses and individuals. Users can experience 60 minutes of free transcription and process up to 5,000 characters of text without requiring a credit card, with the flexibility to cancel anytime. Krybe's services are tailored to maintain a unique brand voice across platforms, facilitating narration, automation, and personalization. The platform aims to streamline workflows, enhance productivity, and enable effortless scaling for its users. Krybe's voice agents are designed to integrate seamlessly with existing systems, functioning like real human assistants to automate business processes. Listen to a real customer service interaction handled seamlessly by our AI voice agent. Effortlessly convert speech to text in real-time, ensuring you never miss a detail while staying fully engaged in discussions.Starting Price: $13 per month -
38
FonadaLabs
FonadaLabs
FonadaLabs is a voice AI platform that provides enterprise-grade infrastructure and APIs for building voice agents on Indian telephony networks. The platform offers a complete voice pipeline that includes telephony hosting, noise cancellation, speech recognition, voice models, and text-to-speech capabilities within a unified API environment. FonadaLabs supports over 23 Indian languages with speech recognition optimized for regional accents and telephony use cases. The platform enables real-time voice streaming with ultra-low latency, enterprise security, and India-based data residency for compliance and sovereignty requirements. Businesses can also leverage specialized voice agent language models, tool-calling support, and natural-sounding Indian voice generation for customer interactions and automation.Starting Price: $5 -
39
Aethex
Aethex
AethexAI is the voice AI stack for emerging markets, built for end-to-end voice agents localized for your market. It brings together infrastructure, models, and deployment in one environment, with proprietary Kora 1 models trained on real conversational speech and human-labeled data across emerging markets. The Kora 1 Engine is designed for real speech, native tool calling, workflow-aware routing, dedicated infrastructure, dialect-aware interactions, and sub-500ms turn-taking. Teams can design, deploy, and manage voice agents that handle calls, messages, and workflows across support, sales, onboarding, and collections, integrated with the systems they already run. It moves from hello to resolution, with agents that can read and write data, trigger actions, and close loops inside existing systems rather than alongside them. Agent Studio lets users design conversation flows, set guardrails, configure personas, and build inbound or outbound agents with no code required.Starting Price: $3 per month -
40
Bond
Bond
Bond is the AI Chief of Staff every founder deserves. It connects to your tools, learns how your company works, and tells you your highest-leverage move. Built for CEOs, founders, and busy executives, BOND gives you a real-time pulse on your company without more meetings, manual updates, or scattered searches across Slack, email, calendar, Notion, Linear, and other tools. It helps leaders understand what needs attention now, what can wait, and where their time should go. Bond preps meetings, reorganizes calendars, protects time for the work that matters most, and turns company noise into a clear daily operating view. Its daily briefing pipeline runs specialized AI agents in parallel to extract todos, summarize updates, prepare meetings, track objectives, and surface what matters. BondBot, its conversational AI agent, orchestrates dozens of skill sets across multiple platforms, searching Slack threads, triaging Linear issues, drafting replies, managing todos, and more.Starting Price: $99 per month -
41
Rubil
Rubil
Voice dictation for Gmail, Slack, Notion + 20 apps. Auto-formats your speech. Audio never stored. 1000 words free daily. Voice dictation for every web app you use. Speak naturally — Rubil formats your speech into clean, ready-to-send text. Properly structured emails. Concise chat messages. Clean document prose. Works across 20+ web apps where knowledge workers spend their day. No cleanup. No rewrites. No copy-paste. No post-editing. Audio is processed instantly through secure transcription and never stored. No transcript history. No voice files on our servers. Your glossary is encrypted both on your device and in the cloud. Teach Rubil your world. Add names, acronyms, and jargon once. Rubil applies them every time you dictate. Voice dictation and voice typing in one click: 1) Hit the mic 2) Speak naturally — ramble, self-correct, think out loud 3) Rubil formats your speech and drops it right in. Done. Free: 1,000 words/day. Pro: $9/mo for unlimited.Starting Price: $9/month -
42
Concierge AI
Concierge AI
Concierge AI is an advanced AI-powered assistant designed to bridge the gap between artificial intelligence and personalized workflow automation. Unlike traditional AI assistants that provide generic responses, Concierge AI connects directly to popular SaaS applications like Gmail, Slack, Notion, Jira, Linear, Attio, and HubSpot, enabling real-time data retrieval and task execution. Users can connect their favorite apps effortlessly, allowing the AI to read and write data in real time, ensuring a smooth workflow without switching between platforms. Concierge AI provides access to top-tier AI models such as GPT, Claude, Grok, and DeepSeek under a single subscription, eliminating the hassle of managing multiple AI services. Whether it’s writing a PRD in a specific format or drafting a sales email in a unique voice, Concierge AI adapts to user preferences, making automation more personalized and efficient. Users can ask Concierge AI to analyze their past communications.Starting Price: $20 per month -
43
Jubilee Voice
Jubilee Voice
Jubilee Voice offers AI-powered voice agents designed to ensure you never miss a call while optimizing costs. These AI agents operate 24/7, scale instantly, and continuously learn to improve performance. Unlike traditional IVR systems, Jubilee Voice’s AI VoiceBot understands caller intent and gets straight to the point without forcing users through lengthy menus. The platform integrates seamlessly with backend systems like Google Calendar and CRMs, automating meeting scheduling and data management. It personalizes interactions by recognizing callers and their previous history, creating a more engaging experience. With features like human override and post-call sentiment analysis, Jubilee Voice combines AI efficiency with empathetic customer service.Starting Price: $0 -
44
Intervo.ai
Intervo.ai
Intervo is an open source, enterprise-grade voice and chat AI agent platform designed to automate real-time customer interactions across voice and text channels. It allows businesses to build, train, and deploy custom agents in minutes without code; you define the agent’s purpose, upload domain knowledge (documents, files), choose a voice engine (e.g., ElevenLabs, Azure), and publish it to embedded channels. Its agents support use cases like lead qualification, customer support, AI receptionist/scheduling, interactive product assistance, and internal help agents (for HR, IT, etc.). They can integrate with telephony via Twilio, connect to multiple LLM backends (OpenAI, Claude, Gemini), orchestrate AI workflows, and embed on websites as widgets. It emphasizes scalability, compliance, and flexibility, letting organizations embed context-aware conversational agents that understand complex queries, route calls, and interact via speech or chat.Starting Price: $10 per month -
45
Runbear
Runbear
Runbear lets teams build shared AI teammates that work inside Slack and Microsoft Teams. Each teammate can join conversations, read company context, and use 2,000+ connected tools to finish workflows without sending people to another app. Teams use Runbear to draft CRM updates, route support requests to the right owner, prepare meeting briefs, answer from approved company knowledge, create Jira or Linear tickets, and automate onboarding follow-ups. Shared agents support per-user authorization, so actions in systems such as HubSpot, Salesforce, Zendesk, Google Drive, Notion, and Confluence run with the right permissions. Runbear connects company knowledge with business tools while keeping each workflow in the conversation where work begins. Teams can configure and deploy agents in minutes without code. Runbear is SOC 2 Type II certified and provides audit logs for shared-agent activity.Starting Price: $79 per month -
46
OpenAI Realtime API
OpenAI
The OpenAI Realtime API is a newly introduced API, announced in 2024, that allows developers to create applications that facilitate real-time, low-latency interactions, such as speech-to-speech conversations. This API is designed for use cases like customer support agents, AI voice assistants, and language learning apps. Unlike previous implementations that required multiple models for speech recognition and text-to-speech conversion, the Realtime API handles these processes seamlessly in one call, enabling applications to handle voice interactions much faster and with more natural flow. -
47
Gemini Audio
Google
Gemini Audio is a set of advanced real-time audio models built on Gemini's architecture, designed to enable natural, fluid voice interaction and expressive audio generation through simple language prompts. It supports conversational experiences where users can speak, listen, and interact with AI in a seamless loop, combining understanding, reasoning, and response generation in audio form. It is capable of both analyzing and generating audio, allowing applications such as speech-to-text transcription, translation, speaker identification, emotion detection, and detailed audio content analysis. They are optimized for low-latency, real-time use cases, making them suitable for live assistants, voice agents, and interactive systems that require continuous, multi-turn dialogue. Gemini Audio also integrates advanced capabilities like function calling, enabling the model to trigger external tools and incorporate real-time data into responses.Starting Price: Free -
48
Coldi
Coldi AI
Coldi is a brand-tuned AI voice agent platform designed to handle sales, support, outreach, and customer engagement at scale. It provides businesses with realistic AI talkers trained to sound human, respond intelligently, and deliver consistent performance across every call. The platform manages everything from voice selection to implementation, ensuring seamless deployment without technical friction. Coldi integrates with major business tools like Twilio, HubSpot, Slack, Zapier, Google Sheets, and more to fit naturally into existing workflows. With capabilities such as lead qualification, appointment setting, abandoned lead recovery, surveys, and customer service automation, it acts as a fully operational voice team. By delivering natural, persuasive conversations, Coldi helps businesses convert more leads, enhance efficiency, and reduce operational costs.Starting Price: $299/month -
49
Utterly Voice
Utterly Voice
Utterly Voice is a highly customizable voice dictation and computer control application designed for a completely hands-free computing experience. It allows users to type text, edit content, press keyboard shortcuts, manage windows, scroll content, control the mouse, and create macros using only their voice. Compatible with Windows 10 and 11, Utterly Voice supports English language input, with plans for additional language support in the future. The application offers multiple speech recognizers and models to choose from, including Vosk, Microsoft Azure, Deepgram, Google Cloud Speech-to-Text V1, and Whisper. Users can easily type individual letters, alphanumerics, or code, and benefit from powerful customization abilities using text configuration files. Advanced mouse control methods, configurable voice commands, and control over speech recognition bias enhance the user experience.Starting Price: Free -
50
SkipCalls
SkipCalls
SkipCalls is a comprehensive AI voice agent platform that revolutionizes phone communication for both businesses and consumers. For B2B clients, it provides 24/7 AI phone agents with deep integrations including CRM systems (Salesforce, HubSpot), calendar platforms (Google Calendar, Outlook), and helpdesk solutions. The platform offers advanced voice AI capabilities with natural language processing, real-time transcription and analytics, customizable AI personas tailored to brand voice. For B2C users, SkipCalls acts as an AI-powered voicemail and outbound calling assistant, eliminating phone anxiety by handling appointment bookings, call screening, spam filtering, and providing instant call summaries. The platform supports webhooks, REST API, and Model Context Protocol (MCP) for seamless workflow integration, making it ideal for healthcare providers, legal practices, retail businesses, and service providers who need to automate routine calls.Starting Price: $3.99