Showing 65 open source projects for "token system"

View related business solutions
  • $300 Free Credits for Your Google Cloud Projects Icon
    $300 Free Credits for Your Google Cloud Projects

    Start building on Google Cloud with $300 in free credits. No commitment, no credit card required until you're ready to scale.

    Launch your next project with $300 in free Google Cloud credits—no strings attached. Test, build, and deploy without risk. Use your credits across the entire Google Cloud platform to find what works best for your needs. After your credits are used, continue with always-free tier services. Only pay when you're ready to scale. Sign up in minutes and start exploring.
    Start Free Trial
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • 1
    Qwen-of-Death
    Qwen of Death is a desktop coding assistant with GUI powered by the Qwen API. Users supply their own API key — all billing is handled directly with Openrouter.ai.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Gemini of Death
    Gemini of Death is a desktop coding assistant with GUI powered by the Googles Gemini API. Users supply their own API key — all billing is handled directly with Google.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 3
    GPT of Death
    GPT of Death is a desktop coding assistant with GUI powered by the OpenAI API. Users supply their own API key — all billing is handled directly with OpenAI.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 4
    Grok of Death
    Grok of Death is a desktop coding assistant with GUI powered by the Grok API. Users supply their own API key — all billing is handled directly with xAI.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • 5
    Claude of Death
    Claude of Death is a desktop coding assistant with GUI powered by the Anthropic Claude API. Users supply their own API key — all billing is handled directly with Anthropic.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    HasMCP

    HasMCP

    Convert API into MCP Server in seconds

    HasMCP empowers AI development by seamlessly connecting your existing APIs to Large Language Models. Its Automated OpenAPI Mapping instantly translates API documentation into LLM-usable tools, eliminating manual coding. Security is paramount, with Native MCP Elicitation Auth managing complex authentication flows like OAuth2, ensuring user credentials are never exposed. To enhance efficiency, Context Window Optimization intelligently prunes API responses using JMESPath and Goja (JS) logic,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Extended Dreambooth How-To Guides

    Extended Dreambooth How-To Guides

    Implementation of Dreambooth

    ...The project adapts and expands upon earlier DreamBooth research by providing practical scripts, notebooks, and workflows that allow users to train personalized models on local machines, cloud environments, or platforms such as Google Colab. It focuses heavily on usability, offering detailed guides for different setups while still exposing advanced configuration options for experienced users. The system allows users to bind a unique token to a subject, which can later be used in prompts to generate consistent and recognizable outputs across different contexts. It also supports captioning, multi-subject training, and regularization techniques to improve generalization and avoid overfitting.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Mixtral offloading

    Mixtral offloading

    Run Mixtral-8x7B models in Colab or consumer desktops

    ...The project implements techniques that allow model components to be dynamically moved between CPU memory and GPU memory during inference, significantly reducing the amount of GPU VRAM required to run the model. This approach takes advantage of the sparse activation properties of mixture-of-experts architectures, where only a subset of expert networks are used for each token during generation. By selectively loading and caching the required experts, the system avoids keeping the entire model in GPU memory at once. The repository includes notebooks and code examples that demonstrate how to run large language models on consumer hardware such as personal GPUs or cloud notebook environments.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    gpu_poor

    gpu_poor

    Calculate token/s & GPU memory requirement for any LLM

    gpu_poor is an open-source tool designed to help developers determine whether their hardware is capable of running a specific large language model and to estimate the performance they can expect from it. The project focuses on calculating GPU memory requirements and predicted inference speed for different models, hardware configurations, and quantization strategies. By analyzing factors such as model size, context length, batch size, and GPU specifications, the system estimates how much VRAM...
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 10
    MiMo-V2.5-Pro

    MiMo-V2.5-Pro

    Flagship MoE model for long-context agents and complex coding

    ...It features approximately 1.02 trillion total parameters with 42B activated per inference, balancing extreme capability with efficient execution. The model supports a 1 million token context window, enabling it to maintain coherence across long workflows involving thousands of tool calls and multi-step reasoning chains. Architecturally, it uses a hybrid attention system combining Sliding Window Attention and Global Attention to significantly reduce memory usage while preserving long-context performance. It also integrates multi-token prediction modules that accelerate inference and improve reinforcement learning efficiency. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    ...It also integrates advanced components such as multi-token prediction modules and specialized vision and audio encoders, making it well-suited for autonomous agents and software development.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Gemma 4 12B

    Gemma 4 12B

    Unified multimodal Gemma model for local coding and reasoning

    ...It supports text, image, audio, and video inputs with text output, making it useful for transcription, image understanding, video analysis, coding, and agentic workflows. The model has 11.95B parameters, 48 layers, a 256K-token context window, and support for over 140 languages. It also includes configurable thinking modes, native system prompt support, function calling, and strong benchmark performance for its size. It is optimized for consumer GPUs, workstations, and streamlined local deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Ministral 3 3B Reasoning 2512

    Ministral 3 3B Reasoning 2512

    Compact 3B-param multimodal model for efficient on-device reasoning

    ...Despite its modest size, the model is designed for edge deployment and can run locally, fitting in ~16 GB of VRAM in BF16 or under 8 GB of RAM/VRAM when quantized. It supports dozens of languages, allowing it to function across global and multilingual contexts. The model retains strong system-prompt adherence, supports function-calling with structured JSON output, and offers a large 256k token context window for extended context reasoning.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Rampart

    Rampart

    Lightweight on-device model for private AI text redaction

    ...It works alongside a deterministic recognizer that handles structured identifiers such as phone numbers and IDs, forming a defense-in-depth client-side redaction system. Rampart supports English, Spanish, French, German, Italian, Portuguese, and Dutch, and is designed to run efficiently even on older devices, making it suitable for browsers, mobile and applications.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Laguna M.1

    Laguna M.1

    Flagship Poolside model for agentic coding and software engineering

    Laguna M.1 is Poolside’s flagship Mixture-of-Experts model built specifically for agentic coding, software engineering, and long-horizon autonomous workflows. It contains approximately 225.8B total parameters with 23.4B activated per token, making it substantially larger and more capable than Laguna XS.2 while maintaining efficient inference through sparse activation. Trained from scratch on roughly 30 trillion tokens using Poolside’s in-house “Model Factory” pipeline, the model focuses on...
    Downloads: 0 This Week
    Last Update:
    See Project
Auth0 Logo