22 Integrations with Cerebras

View a list of Cerebras integrations and software that integrates with Cerebras below. Compare the best Cerebras integrations as well as features, ratings, user reviews, and pricing of software that integrates with Cerebras. Here are the current Cerebras integrations in 2026:

  • 1
    StackAI

    StackAI

    StackAI

    StackAI is an enterprise AI automation platform to build end-to-end internal tools and processes with AI agents in a fully compliant and secure way. Designed for large, regulated organizations, it enables teams to automate complex workflows across operations, compliance, finance, IT, and support without heavy engineering. With StackAI you can: • Connect knowledge bases (SharePoint, Confluence, Notion, Google Drive, databases) with versioning, citations, and access controls • Publish AI agents as chat assistants, advanced forms, or APIs integrated into Slack, Teams, Salesforce, HubSpot, or ServiceNow • Govern usage with enterprise security: SSO (Okta, Azure AD, Google), RBAC, audit logs, PII masking, data residency, and cost controls • Route across OpenAI, Anthropic, Google, or local LLMs with guardrails, evaluations, and testing • Deploy in multi-tenant cloud, dedicated cloud, private cloud, or on-premise
    Leader badge
    Starting Price: $0
    View Software
    Visit Website
  • 2
    MachinesFluent

    MachinesFluent

    MachinesFluent

    MachinesFluent is an extremely flexible AI‑powered dictation app. It lets you dictate in any app, work offline or online, and turn speech into raw text, polished writing, summaries, translations, replies, documentation, structured notes, or any custom format you define. You can: - ask web‑search questions by voice - process copied text - analyze clipboard images - transcribe existing audio or video files MachinesFluent gives you control over the engine behind each task. Features include: - offline dictation for privacy - cloud speech for convenience - local AI - cloud AI - direct OpenAI account sign‑in - custom prompts - per‑prompt model choices - vocabulary dictionaries - voice snippets - local history - configurable hotkeys - app or website‑aware dictation styles It is for users who want dictation to be fast, private when needed, AI‑powered when useful, and flexible enough to match how they actually work.
    Starting Price: $9/month/user
  • 3
    Kimi K2.7 Code

    Kimi K2.7 Code

    Moonshot AI

    Kimi K2.7 Code is an open-source, coding-focused agentic AI model developed by Moonshot AI for long-horizon software engineering tasks. It is designed to improve coding performance, agent workflows, and real-world development assistance compared with earlier Kimi K2 versions. The model supports a 256K context window, making it useful for working with large codebases, long technical documents, and complex multi-step programming tasks. Kimi K2.7 Code is available through Kimi Code and API access, with OpenAI- and Anthropic-compatible options for easier integration into developer workflows. It is also listed on Hugging Face and supports deployment through inference engines such as vLLM, SGLang, and KTransformers. With improved agentic capabilities, long-context support, and reduced thinking-token usage compared with K2.6, Kimi K2.7 Code gives developers a flexible open-source option for AI-assisted coding.
    Starting Price: Free
  • 4
    SWE-1.7

    SWE-1.7

    Cognition

    SWE-1.7 is Cognition’s frontier software engineering model designed to deliver high intelligence at a lower rollout cost. The model is optimized for long-horizon agentic coding tasks, including debugging, feature implementation, codebase exploration, migrations, terminal workflows, and multilingual software engineering. SWE-1.7 was trained from a Kimi K2.7 base using large-scale reinforcement learning improvements across infrastructure, data quality, training stability, self-compaction, and long-running task execution. It is built to explore codebases thoroughly, probe edge cases, identify hidden requirements, and produce more complete end-to-end solutions. The model is available in Devin across web, desktop, and CLI through Cerebras at very high serving speeds. SWE-1.7 is positioned for developers and engineering teams that need cost-efficient frontier-level coding intelligence for complex real-world software work.
    Starting Price: $20/month
  • 5
    Kimi K3

    Kimi K3

    Moonshot AI

    Kimi K3 is Moonshot AI’s most capable model, built for frontier intelligence scenarios such as software engineering, knowledge work, deep reasoning, and multimodal understanding. The model has 2.8 trillion parameters and uses Kimi Delta Attention, a hybrid linear attention mechanism, along with Attention Residuals for long-context performance. Kimi K3 supports a 1 million token context window, making it useful for analyzing large codebases, long documents, complex knowledge bases, and multi-step workflows. It includes native visual understanding for images and videos, with support for structured message formats, base64 image input, uploaded video files, and multimodal reasoning. Developers can use Kimi K3 through an OpenAI-compatible API with support for streaming, structured JSON output, partial mode, custom tools, dynamic tool loading, and automatic context caching.
    Starting Price: $3 per 1M tokens (input)
  • 6
    Kimi K2.6

    Kimi K2.6

    Moonshot AI

    Kimi K2.6 is a next-generation agentic AI model developed by Moonshot AI, designed to push forward real-world execution, coding, and multi-step reasoning beyond earlier K2 and K2.5 versions. It builds on a Mixture-of-Experts architecture and the multimodal, agent-first foundation of the Kimi series, combining language understanding, coding, and tool use into a single system capable of planning and executing complex workflows. It introduces deeper reasoning capabilities and significantly improved agent planning, allowing it to break down tasks, coordinate tools, and handle multi-file or multi-step problems with greater accuracy and efficiency. It supports advanced tool calling with high reliability, enabling integration with external systems such as web search or APIs, and includes built-in validation mechanisms to ensure correct execution formats.
    Starting Price: Free
  • 7
    omp

    omp

    omp

    omp is an open source AI coding agent and development harness that provides developers with a powerful local environment for AI-assisted engineering. It connects AI models directly to IDE capabilities, debugging tools, code execution, language servers, browser automation, memory, and dozens of built-in development tools. It supports more than 40 AI providers while allowing developers to use a single interface across cloud and local language models. omp enhances coding performance with features such as intelligent code editing, parallel subagents, persistent execution environments, integrated debugging, and advanced code review workflows. It also includes collaborative sessions, local memory, workflow automation, browser control, and GitHub integration to streamline complex software development tasks. Built with a native Rust engine and designed for Windows, macOS, and Linux, omp helps developers build, debug, and maintain software.
    Starting Price: Free
  • 8
    MindMac

    MindMac

    MindMac

    MindMac is a native macOS application designed to enhance productivity by integrating seamlessly with ChatGPT and other AI models. It supports multiple AI providers, including OpenAI, Azure OpenAI, Google AI with Gemini, Gemini Enterprise Agent Platform, Anthropic Claude, OpenRouter, Mistral AI, Cohere, Perplexity, OctoAI, and local LLMs via LMStudio, LocalAI, GPT4All, Ollama, and llama.cpp. MindMac offers over 150 built-in prompt templates to facilitate user interaction and allows for extensive customization of OpenAI parameters, appearance, context modes, and keyboard shortcuts. The application features a powerful inline mode, enabling users to generate content or ask questions within any application without switching windows. MindMac ensures privacy by storing API keys securely in the Mac's Keychain and sending data directly to the AI provider without intermediary servers. The app is free to use with basic features, requiring no account for setup.
    Starting Price: $29 one-time payment
  • 9
    Sonar

    Sonar

    Perplexity

    Perplexity has recently introduced an enhanced version of its AI search engine, named Sonar. Built upon the Llama 3.3 70B model, Sonar has undergone additional training to improve the factual accuracy and readability of responses in Perplexity's default search mode. This advancement aims to deliver users more precise and comprehensible answers while maintaining the platform's characteristic efficiency and speed. Sonar also provides real-time, web-wide research and Q&A capabilities, allowing developers to integrate these features into their products through a lightweight, cost-effective, and user-friendly API. The Sonar API supports advanced models like sonar-reasoning-pro and sonar-pro, designed for complex tasks requiring deep understanding and context retention. These models offer detailed answers with an average of twice as many citations as previous versions, enhancing the transparency and reliability of the information provided.
    Starting Price: Free
  • 10
    LiteLLM

    LiteLLM

    LiteLLM

    ​LiteLLM is a versatile platform designed to streamline interactions with over 100 Large Language Models (LLMs) through a unified interface. It offers both a Proxy Server (LLM Gateway) and a Python SDK, enabling developers to integrate various LLMs seamlessly into their applications. The Proxy Server facilitates centralized management, allowing for load balancing, cost tracking across projects, and consistent input/output formatting compatible with OpenAI standards. This setup supports multiple providers. It ensures robust observability by generating unique call IDs for each request, aiding in precise tracking and logging across systems. Developers can leverage pre-defined callbacks to log data using various tools. For enterprise users, LiteLLM offers advanced features like Single Sign-On (SSO), user management, and professional support through dedicated channels like Discord and Slack.
    Starting Price: Free
  • 11
    Axolotl

    Axolotl

    Axolotl

    ​Axolotl is an open source tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures. It enables users to train models, supporting methods like full fine-tuning, LoRA, QLoRA, ReLoRA, and GPTQ. Users can customize configurations using simple YAML files or command-line interface overrides, and load different dataset formats, including custom or pre-tokenized datasets. Axolotl integrates with technologies like xFormers, Flash Attention, Liger kernel, RoPE scaling, and multipacking, and works with single or multiple GPUs via Fully Sharded Data Parallel (FSDP) or DeepSpeed. It can be run locally or on the cloud using Docker and supports logging results and checkpoints to several platforms. It is designed to make fine-tuning AI models friendly, fast, and fun, without sacrificing functionality or scale.
    Starting Price: Free
  • 12
    Mastra AI

    Mastra AI

    Mastra AI

    Mastra is a powerful TypeScript framework for building intelligent AI agents that can execute tasks, access knowledge bases, and maintain memory persistently within workflows. This framework simplifies the process of creating and deploying AI-powered agents by leveraging TypeScript’s capabilities to streamline development. With features like customizable agent instructions, memory, and task orchestration, Mastra provides developers with the tools to build and scale AI agents for various applications, from personal assistants to specialized domain experts.
    Starting Price: Free
  • 13
    Vivgrid

    Vivgrid

    Vivgrid

    Vivgrid is a development platform for AI agents that emphasizes observability, debugging, safety, and global deployment infrastructure. It gives you full visibility into agent behavior, logging prompts, memory fetches, tool usage, and reasoning chains, letting developers trace where things break or deviate. You can test, evaluate, and enforce safety policies (like refusal rules or filters), and incorporate human-in-the-loop checks before going live. Vivgrid supports the orchestration of multi-agent systems with stateful memory, routing tasks dynamically across agent workflows. On the deployment side, it operates a globally distributed inference network to ensure low-latency (sub-50 ms) execution and exposes metrics like latency, cost, and usage in real time. It aims to simplify shipping resilient AI systems by combining debugging, evaluation, safety, and deployment into one stack, so you're not stitching together observability, infrastructure, and orchestration.
    Starting Price: $25 per month
  • 14
    Rebolt.ai

    Rebolt.ai

    Rebolt.ai

    Rebolt is an enterprise-grade AI platform that empowers companies to build custom applications and intelligent agents simply by speaking with AI. The platform integrates seamlessly with corporate tools such as OneDrive, SharePoint, Salesforce, Slack, and custom APIs, offering built-in infrastructure like databases, file storage, scheduling (cron jobs), audit logs, and distinct staging and production environments for deployment. Users can create apps and agents without writing API key code by simply describing what they need in natural language, while maintaining enterprise-class security, permissions mapping (e.g., via Azure groups), and role-based access. Rebolt is designed for building operational workflows, internal tools, and automation that connect to existing company data and services, enabling non-technical users (or low-code teams) to rapidly assemble solutions and replace spreadsheets, manual processes, and fragmented SaaS stacks.
    Starting Price: $25 per month
  • 15
    Zo Computer

    Zo Computer

    Zo Computer

    Zo Computer is an always-on AI companion designed to act like your own personal cloud computer. It works 24/7 to schedule meetings, clean your inbox, organize files, and run tasks while you’re away. Users can interact with Zo through its app or simply by texting it commands. Built on a powerful Linux server, Zo gives you full control to host files, build automations, and run projects effortlessly. It supports deep research, web browsing, reminders, and data organization in one unified environment. Zo combines AI, code, and compute into a single system you own. It’s built to help you get real work done, not just chat.
    Starting Price: $18/month
  • 16
    Anuma

    Anuma

    Anuma

    Anuma is a privacy-first, multi-model AI platform that unifies access to leading proprietary and open-source AI systems within a single interface while giving users full ownership and control over their data. It allows users to interact with models such as ChatGPT, Claude, Gemini, Grok, and open source alternatives like DeepSeek or Qwen without switching tools or losing context, enabling seamless workflows across different AI engines. At its core is a Private Memory Layer that stores user preferences, conversation history, and context in an encrypted, user-controlled environment, ensuring that sensitive data is not accessible to providers or stored centrally. This memory persists across sessions and models, allowing users to continue tasks without re-explaining information and maintaining continuity in complex workflows. It supports comparing multiple models simultaneously, building custom mini-apps and automations without code.
    Starting Price: $9.99 per month
  • 17
    Pi Agent
    Pi is a minimal terminal coding harness built to adapt to developer workflows instead of forcing developers to adapt to it. It ships with powerful defaults, but stays intentionally small and aggressively extensible, letting users customize Pi with extensions, skills, prompt templates, themes, and shareable packages from npm or git. If a team needs a command, tool, provider, workflow, or UI tweak, they can ask Pi to build it, manipulate it in place, reload, and keep going. Pi supports interactive, print/JSON, RPC, and SDK modes, making it usable as a full terminal UI, a scriptable command, a JSON event stream, or an embeddable agent harness. It works with 15+ providers and hundreds of models, including Anthropic, OpenAI, Google, Azure, Bedrock, Mistral, Groq, Cerebras, xAI, Hugging Face, Kimi For Coding, MiniMax, OpenRouter, Ollama, and more, with mid-session model switching.
    Starting Price: Free
  • 18
    Workers by Delos
    AI Workers are autonomous agents built for your business; specialized AI workers that act like real coworkers, not chatbots you prompt. They come with their own professional profile, email, phone number, Slack and Teams presence, initiative, and the ability to work 24/7 without waiting to be asked. Instead of telling them exactly how to complete every step, you set the goal, and they build the workflow, whether that means daily reports, weekly follow-ups, CRM updates, client communication, research, content, finance tasks, HR coordination, design work, development support, or other recurring business operations. It includes specialized AI Workers across marketing, development, design, HR, finance, and other business functions, with each worker designed around a clear role and practical use cases. AI Workers can connect to more than 3,000 tools, including Slack, Microsoft Teams, Gmail, Notion, HubSpot, Salesforce, and other business apps.
    Starting Price: $30 per month
  • 19
    flo2

    flo2

    Data Products LLP

    flo2 is an LLM gateway and router that provides access to major AI model providers (OpenAI, Anthropic, Groq, Cerebras, DeepInfra) through one unified, OpenAI-compatible API. Smart routing picks the cheapest or fastest model per request. Automatic fallback keeps applications running when a provider goes down. Racing mode runs requests across providers in parallel. Full cost accounting per request, per model, per project. Developers use their own provider keys via flo2.com — RapidAPI's testing tier includes free tokens for evaluation.
    Starting Price: 0
  • 20
    Mercury 2

    Mercury 2

    Inception

    Mercury 2 is the first reasoning model fast enough to pick up the phone, a reasoning diffusion language model built for real-time voice agents. Instead of making callers wait through seconds of dead air while an autoregressive model generates thinking tokens one by one, Mercury 2 uses a diffusion large language model architecture to generate tokens in parallel, decoding 1000+ tokens per second on standard NVIDIA GPUs. That speed is fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation, reducing the cost of reasoning from seconds of silence to roughly 300 milliseconds. Mercury models work by corrupting clean text into noise, then training a standard Transformer to reverse the process and predict clean text across all positions simultaneously. Because each denoising pass touches many tokens, generation uses the GPU more efficiently than one-token-at-a-time decoding, making custom-silicon-like speed possible on NVIDIA H100s.
  • 21
    Cirrascale

    Cirrascale

    Cirrascale

    Our high-throughput storage systems can serve millions of small, random files to GPU-based training servers accelerating overall training times. We offer high-bandwidth, low-latency networks for connecting distributed training servers as well as transporting data between storage and servers. Other cloud providers squeeze you with extra fees and charges to get your data out of their storage clouds, and those can add up fast. We consider ourselves an extension of your team. We work with you to set up scheduling services, help with best practices, and provide superior support. Workflows can vary from company to company. Cirrascale works to ensure you get the right solution for your needs to get you the best results. Cirrascale is the only provider that works with you to tailor your cloud instances to increase performance, remove bottlenecks, and optimize your workflow. Cloud-based solutions to accelerate your training, simulation, and re-simulation time.
    Starting Price: $2.49 per hour
  • 22
    Witsy

    Witsy

    Witsy

    Witsy is a desktop application that provides access to all generative AI models from top AI providers. It is a one-stop solution for all your generative AI needs. Witsy is a BYOK (Bring Your Own Keys) AI application, meaning you need to have API keys for the LLM providers you want to use. Alternatively, you can use Ollama to run models locally on your machine for free and use them in Witsy. Witsy itself does not collect or process any personal data. All your data is kept on your computer and never leaves it. Witsy does not use cookies or any other tracking mechanisms. All features of Witsy are accessible through keyboard shortcuts. You can trigger the chat or scratchpad, execute commands, and more. You can also customize the shortcuts to fit your needs. You can also talk to the AI model in the scratchpad for the ultimate document creation experience. Work with Witsy as if you were collaborating with a peer.
  • Previous
  • You're on page 1
  • Next