Showing 4 open source projects for "masmt runtime engine"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Get Started Free
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Try Free
  • 1
    Colibrì

    Colibrì

    Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine

    Colibri is a compact inference engine designed to run the 744-billion-parameter GLM-5.2 mixture-of-experts model on consumer hardware. It keeps the dense portion of the quantized model in memory while streaming routed experts from a large disk-based store as they are needed. The runtime is implemented in pure C, requires no Python or BLAS during inference, and can operate without a GPU.
    Downloads: 9 This Week
    Last Update:
    See Project
  • 2
    Kimi K3 in C

    Kimi K3 in C

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU

    Kimi K3 in C is a portable C99 inference engine built to run the 2.78-trillion-parameter Kimi K3 model on CPUs without BLAS, machine-learning frameworks, or GPUs. It demonstrates inference from a roughly 1.56 TB checkpoint with measured memory use as low as 8.24 GB. The runtime streams model trunk layers and routed experts from disk instead of keeping all weights resident in memory.
    Downloads: 21 This Week
    Last Update:
    See Project
  • 3
    WASTE

    WASTE

    Run the full 2.78-trillion-parameter Kimi K3 model

    ...Its main target is the full 2.78-trillion-parameter Kimi K3 model, including multimodal image input. The engine has no third-party runtime dependency on its CPU inference path and exposes both a CLI and C library. An optional server provides an OpenAI-compatible chat API with streaming, tools, structured output, and image support. Validation compares layers, logits, vision output, and prompt rendering against reference implementations.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 4
    Proximus for Ryzen AI

    Proximus for Ryzen AI

    Runtime extension of Proximus enabling Deployment on AMD Ryzen™ AI

    This project extends the Proximus development environment to support deployment of AI workloads on next-generation AMD Ryzen™ AI processors, such as the Ryzen™ AI 7 PRO 7840U featured in the Lenovo ThinkPad T14s Gen 4 ,one of the first true AI PCs with an onboard Neural Processing Unit (NPU) capable of 16 TOPS (trillion operations per second). Originally designed for use with Windows 11 Pro, this runtime was further enhanced to work under Linux environments, allowing developers and researchers to fully utilize the AMD AI Engine across both platforms. This cross-platform support is a major innovation, enabling AI workload portability, integration into CI environments, and deployment into Linux-based research and production pipelines.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Train ML Models With SQL You Already Know Icon
    Train ML Models With SQL You Already Know

    BigQuery automates data prep, analysis, and predictions with built-in AI assistance.

    Build and deploy ML models using familiar SQL. Automate data prep with built-in Gemini. Query 1 TB and store 10 GB free monthly.
    Try Free
  • Previous
  • You're on page 1
  • Next