Showing 4 open source projects for "inference engine"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    uzu

    uzu

    A high-performance inference engine for AI models

    uzu is a high-performance inference engine designed to run artificial intelligence models efficiently on Apple Silicon hardware. Written primarily in Rust and leveraging Apple’s Metal framework, the project focuses on maximizing performance when executing large language models and other AI workloads on devices such as Mac computers with M-series chips. The engine implements a hybrid architecture in which model layers can be executed either as custom GPU kernels or through Apple’s MPSGraph API, allowing it to balance performance and compatibility depending on the workload. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    BaseRT

    BaseRT

    Fastest LLM inference runtime for Apple Silicon

    ...Its server implements OpenAI-compatible chat, completion, embedding, transcription, tool-calling, and multimodal endpoints. The custom .base format supports affine quantization from Q2 through Q8, optional AWQ calibration, and signed model bundles. Stable C interfaces connect the engine with Python, Node.js, Rust, and Swift applications. The repository contains the open CLI, format specifications, bindings, documentation, and benchmarks, while the prebuilt inference engine uses a separate license.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    mistral.rs

    mistral.rs

    Fast, flexible LLM inference

    mistral.rs is a fast and flexible LLM inference engine implemented in Rust, designed to run and serve modern language models with an emphasis on performance and practical deployment. It provides multiple entry points for developers, including a CLI for running models locally and an HTTP server that exposes an OpenAI-compatible API surface for easy integration with existing clients.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 4
    Ante

    Ante

    Ghost in your shell. Ante is a self-contained agent harness

    ...It ships as a roughly 15 MB binary with no external runtime dependencies. Users can work through an interactive terminal interface, headless commands, a server protocol, or Slack and Discord gateways. A built-in inference engine can run GGUF models entirely offline without an account, API key, or internet connection. Ante also supports more than a dozen hosted providers and can switch between commercial, open-weight, and local models. Multi-agent orchestration lets it spawn and coordinate specialized subagents for larger software tasks. Skills, MCP integrations, persistent memory, resumable sessions, and public benchmark evaluation extend it into a lightweight general-purpose agent harness.
    Downloads: 7 This Week
    Last Update:
    See Project
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • Previous
  • You're on page 1
  • Next