Best Artificial Intelligence Software for LiveKit

Compare the Top Artificial Intelligence Software that integrates with LiveKit as of July 2026

This a list of Artificial Intelligence software that integrates with LiveKit. Use the filters on the left to add additional filters for products that have integrations with LiveKit. View the products that work with LiveKit in the table below.

What is Artificial Intelligence Software for LiveKit?

Artificial Intelligence (AI) software is computer technology designed to simulate human intelligence. It can be used to perform tasks that require cognitive abilities, such as problem-solving, data analysis, visual perception and language translation. AI applications range from voice recognition and virtual assistants to autonomous vehicles and medical diagnostics. Compare and read user reviews of the best Artificial Intelligence software for LiveKit currently available using the table below. This list is updated regularly.

  • 1
    Google Cloud Platform
    Google Cloud Platform provides an extensive suite of Artificial Intelligence (AI) and machine learning tools designed to streamline data analysis. GCP offers pre-trained models and APIs like Vision AI, Natural Language, and AutoML that allow businesses to easily incorporate AI into their applications without requiring deep expertise in the field. Additionally, new customers receive $300 in free credits to run, test, and deploy workloads, enabling them to explore AI capabilities on the platform and implement advanced machine learning solutions at no initial cost. GCP’s AI tools also integrate seamlessly with other services, creating end-to-end machine learning pipelines from data processing to model deployment. Furthermore, these tools are designed to be highly scalable, allowing companies to experiment with AI and grow their AI-powered solutions as their needs expand. With these resources, businesses can quickly leverage AI for various tasks, from predictive analytics to automation.
    Leader badge
    Starting Price: Free ($300 in free credits)
    View Software
    Visit Website
  • 2
    Speechmatics

    Speechmatics

    Speechmatics

    Best-in-Market Speech-to-Text & Voice AI for Enterprises. Speechmatics delivers industry-leading Speech-to-Text and Voice AI for enterprises needing unrivaled accuracy, security, and flexibility. Our enterprise-grade APIs provide real-time and batch transcription with exceptional precision—across the widest range of languages, dialects, and accents. Powered by Foundational Speech Technology, Speechmatics supports mission-critical voice applications in media, contact centers, finance, healthcare, and more. With on-prem, cloud, and hybrid deployment, businesses maintain full control over data security while unlocking voice insights. Trusted by global leaders, Speechmatics is the top choice for best-in-class transcription and voice intelligence. 🔹 Unmatched Accuracy – Superior transcription across languages & accents 🔹 Flexible Deployment – Cloud, on-prem, and hybrid 🔹 Enterprise-Grade Security – Full data control 🔹 Real-Time & Batch Processing – Scalable transcription
    Starting Price: $0 per month
  • 3
    OpenAI

    OpenAI

    OpenAI

    OpenAI’s mission is to ensure that artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work—benefits all of humanity. We will attempt to directly build safe and beneficial AGI, but will also consider our mission fulfilled if our work aids others to achieve this outcome. Apply our API to any language task — semantic search, summarization, sentiment analysis, content generation, translation, and more — with only a few examples or by specifying your task in English. One simple integration gives you access to our constantly-improving AI technology. Explore how you integrate with the API with these sample completions.
  • 4
    Character.AI

    Character.AI

    Character.AI

    Character.AI is bringing to life the science-fiction dream of open-ended conversations and collaborations with computers. We are building the next generation of dialog agents; with a long-tail of applications spanning entertainment, education, general question-answering and others. Our dialog agents are powered by our own proprietary technology based on large language models, built and trained from the ground up with conversation in mind. The Character.AI beta is based on neural language models. A supercomputer reads huge amounts of text and learns to hallucinate what words might come next in any given situation. Models like these have many uses including auto-complete and machine translation. At Character.AI, you collaborate with the computer to write a dialog - you write one character's lines, and the computer creates the other character's lines, giving you the illusion that you are talking with the other character.
  • 5
    ai-coustics

    ai-coustics

    ai-coustics

    ai-coustics is a Berlin-based startup building the audio intelligence layer for Voice AI. Founded by researchers in audio, acoustics, and machine learning, the company focuses on the fundamental reliability problem that causes voice systems to fail outside controlled environments. Rather than competing with ASR, LLMs, or TTS, ai-coustics makes them reliable. Its SDK and model infrastructure sit between real-world sound and machine understanding, conditioning raw audio into stable, machine-ready input optimized for downstream behavior. The company’s Quail model family delivers real-time speech enhancement, speaker isolation, and voice activity detection designed specifically for production of Voice AI. ai-coustics powers voice agents, transcription pipelines, and telephony systems, and is natively integrated in LiveKit and Pipecat. Its mission is to make audio input reliable and measurable, so voice systems can operate with confidence where real people actually speak.
    Starting Price: $149 / month
  • 6
    Rime

    Rime

    Rime

    Rime is a next-generation voice AI platform that delivers ultra-natural, emotionally aware text-to-speech technology, enabling enterprises and startups to build applications that convert, retain, and sell. With sub-200ms latency on the cloud (and <100ms on-prem), plus fine-grained voice controls and pronunciation accuracy, Rime is redefining how businesses engage with customers through voice. Founded in 2022 by experts in linguistics and machine learning, Rime combines deep linguistic expertise with advanced AI to create voices that reflect the richness and diversity of human speech. Our proprietary dataset comprises real conversations across various demographics, accents, and languages, ensuring authentic and relatable voice outputs. Rime's technology includes models like Mist and Arcana, which offer features such as paralinguistic expressions and the ability to generate new voices dynamically.
    Starting Price: $5 per month
  • 7
    Gladia

    Gladia

    Gladia

    Gladia is a speech-to-text platform built for production, turning raw audio into structured outputs that power real workflows like meeting summaries, CRM enrichment, contact center QA, and real-time voice assistants. With support for 99+ languages and the ability to handle messy real-world audio—overlapping speakers, accents, code-switching, domain-specific terminology—Gladia is designed for the complexity of actual conversations, not clean studio recordings.
    Starting Price: 10 hours free
  • 8
    EffectsSDK

    EffectsSDK

    EffectsSDK

    EffectsSDK is a cross-platform AI-powered video enhancement SDK that enables developers and businesses to integrate real-time video effects and webcam enhancement features into their applications. Designed for video conferencing platforms, telehealth systems, streaming tools, educational platforms, presentation services, and communication software, EffectsSDK provides advanced AI-based effects such as background blur, virtual background replacement, intelligent auto-framing, skin smoothing, text/graphic overlays, noise suppression, studio sound, and voice changer. The SDK operates in real time and is optimized for performance across Web, Windows, macOS, iOS, Android, and Linux environments, supporting technologies such as DirectX, OpenGL, Metal, OpenVINO, WinML, and CoreML with GPU acceleration. Developers can integrate the SDK into applications using C++, Objective-C, Kotlin, JavaScript, and WebAssembly-based wrappers depending on platform requirements.
    Starting Price: $50/month
  • 9
    Workers by Delos
    AI Workers are autonomous agents built for your business; specialized AI workers that act like real coworkers, not chatbots you prompt. They come with their own professional profile, email, phone number, Slack and Teams presence, initiative, and the ability to work 24/7 without waiting to be asked. Instead of telling them exactly how to complete every step, you set the goal, and they build the workflow, whether that means daily reports, weekly follow-ups, CRM updates, client communication, research, content, finance tasks, HR coordination, design work, development support, or other recurring business operations. It includes specialized AI Workers across marketing, development, design, HR, finance, and other business functions, with each worker designed around a clear role and practical use cases. AI Workers can connect to more than 3,000 tools, including Slack, Microsoft Teams, Gmail, Notion, HubSpot, Salesforce, and other business apps.
    Starting Price: $30 per month
  • 10
    Inworld TTS
    Inworld TTS is a state-of-the-art text-to-speech platform designed to deliver ultra-realistic, context-aware speech synthesis and precise voice-cloning capabilities at a radically accessible price. The flagship model, TTS-1, is optimized for real-time applications and supports low-latency streaming (first audio chunk in ≈200 ms) as well as multiple languages (including English, Spanish, French, Korean, Chinese, and more). Developers can use instant zero-shot voice cloning (5-15 seconds of audio) or professional fine-tuned cloning, add voice-tags for emotion, style, and non-verbal sounds, and switch languages while preserving voice identity. The larger TTS-1-Max model (in preview) offers even more expressive speech and multilingual strength. The platform supports both API and portal access, streaming or batch mode, and is designed for everything from interactive voice agents and gaming characters to branded audio experiences.
    Starting Price: $0.005 per minute
  • 11
    Operata

    Operata

    Operata

    Operata is an AI-powered CX observability platform built exclusively for cloud contact centers that continuously collects and correlates real-time data from every call, agent environment, network, CCaaS and AI interaction to provide end-to-end visibility into customer and agent experience so teams can understand not just what happened but why it happened and act on it quickly; its features include a unified CX Insights Graph that harmonizes technical, operational and experience signals, CX Copilot and Agent Copilot assistants powered by Tenor AI for natural language querying and on-the-fly recommendations, Customer Journey Trace to visualize complete interaction sequences across multiple platforms, pre-built playbooks and interactive dashboards for proactive insights, readiness testing and assurance tools to benchmark performance, seamless integrations with 50+ CX and voice systems, and an MCP Server to feed observability data into enterprise AI stacks.
    Starting Price: $0.0060 per agent minutes
  • 12
    Mercury 2

    Mercury 2

    Inception

    Mercury 2 is the first reasoning model fast enough to pick up the phone, a reasoning diffusion language model built for real-time voice agents. Instead of making callers wait through seconds of dead air while an autoregressive model generates thinking tokens one by one, Mercury 2 uses a diffusion large language model architecture to generate tokens in parallel, decoding 1000+ tokens per second on standard NVIDIA GPUs. That speed is fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation, reducing the cost of reasoning from seconds of silence to roughly 300 milliseconds. Mercury models work by corrupting clean text into noise, then training a standard Transformer to reverse the process and predict clean text across all positions simultaneously. Because each denoising pass touches many tokens, generation uses the GPU more efficiently than one-token-at-a-time decoding, making custom-silicon-like speed possible on NVIDIA H100s.
  • 13
    Oracle Cloud Infrastructure
    Oracle Cloud Infrastructure supports traditional workloads and delivers modern cloud development tools. It is architected to detect and defend against modern threats, so you can innovate more. Combine low cost with high performance to lower your TCO. Oracle Cloud is a Generation 2 enterprise cloud that delivers powerful compute and networking performance and includes a comprehensive portfolio of infrastructure and platform cloud services. Built from the ground up to meet the needs of mission-critical applications, Oracle Cloud supports all legacy workloads while delivering modern cloud development tools, enabling enterprises to bring their past forward as they build their future. Our Generation 2 Cloud is the only one built to run Oracle Autonomous Database, the industry's first and only self-driving database. Oracle Cloud offers a comprehensive cloud computing portfolio, from application development and business analytics to data management, integration, security, AI & blockchain.
  • 14
    Gemini Live API
    ​The Gemini Live API is a preview feature that enables low-latency, bidirectional voice and video interactions with Gemini. It allows end users to experience natural, human-like voice conversations and provides the ability to interrupt the model's responses using voice commands. The model can process text, audio, and video input, and it can provide text and audio output. New capabilities include two new voices and 30 new languages with configurable output language, configurable image resolutions (66/256 tokens), configurable turn coverage (send all inputs all the time or only when the user is speaking), configurable interruption settings, configurable voice activity detection, new client events for end-of-turn signaling, token counts, a client event for signaling the end of stream, text streaming, configurable session resumption with session data stored on the server for 24 hours, and longer session support with a sliding context window.
  • 15
    Kipps.AI

    Kipps.AI

    Kipps.AI

    Kipps.AI is an enterprise-grade platform for building and deploying AI agents, voice, chat, and WhatsApp that can handle millions of conversations with human-like intelligence and enterprise-scale reliability. It enables organizations to deploy custom agents for lead qualification, booking appointments, customer support, and more, with integrations into CRM systems, telephony platforms, and other business tools. It supports 100 + pre-built integrations such as Salesforce, HubSpot, WhatsApp, Slack, and Zoom; features include detailed analytics (model- and agent-level usage), conversation transcription, real-time call-streaming, sentiment detection, routing to human agents when needed, and enterprise-grade security with SOC 2 Type II, ISO 27001, HIPAA-ready, PCI DSS Level 1, and zero-data-retention options.
  • 16
    HeyGen

    HeyGen

    HeyGen

    Meet HeyGen - The best AI video generation platform for your team. Create AI videos in 3 easy steps: 1. Pick your avatar 2. Input your script 3. Submit to generate videos HeyGen is a video platform that help you create engaging business videos with generative AI, as easily as making PowerPoints for various use cases. Create professional business videos for Marketing & Sales, Training & Onboarding and more! Engage your audience with a more personal and inviting video message. Turn your text into a professional video in minutes, right from your browser. Record & upload your real voice to create a personalized Avatar. Choose from 300+ voices in 40+ popular languages. Combine several scenes into one video. End-to-end videos are as easy as PowerPoint slides. Videos come in 1080P with unlimited downloads. HeyGen AI Studio is a cutting-edge video creation platform that uses advanced AI technology to enable users to produce high-quality, customizable videos with ease.
    Starting Price: $24 per month
  • Previous
  • You're on page 1
  • Next
Monday.com Logo