Showing 52 open source projects for "llama.cpp"

View related business solutions
  • $300 Free Credits for Your Google Cloud Projects Icon
    $300 Free Credits for Your Google Cloud Projects

    Start building on Google Cloud with $300 in free credits. No commitment, no credit card required until you're ready to scale.

    Launch your next project with $300 in free Google Cloud credits—no strings attached. Test, build, and deploy without risk. Use your credits across the entire Google Cloud platform to find what works best for your needs. After your credits are used, continue with always-free tier services. Only pay when you're ready to scale. Sign up in minutes and start exploring.
    Start Free Trial
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • 1
    QodeAssist

    QodeAssist

    QodeAssist is an AI-powered coding assistant plugin for Qt Creator

    ...It provides intelligent code completion, inline refactoring, and conversational chat interfaces that can appear as popups, side panels, or embedded views within the IDE. The plugin supports both local and cloud-based language model providers, including Ollama, llama.cpp, OpenAI, Claude, and others, allowing developers to choose between privacy-focused local setups or more powerful remote models. It includes advanced context handling features such as file attachments, project-specific rules, and automatic synchronization with open files, which improves the relevance of AI responses. The system also enables AI tool calling, allowing the assistant to inspect project files, run searches, and access diagnostics to provide deeper assistance.
    Downloads: 6 This Week
    Last Update:
    See Project
  • 2
    MindWork AI Studio

    MindWork AI Studio

    Independent cross-platform desktop app for local and cloud LLMs

    ...It is built with a strong focus on accessibility and democratization, enabling users to run AI workflows even on low-cost hardware while maintaining flexibility in choosing providers such as OpenAI, Gemini, Anthropic, and self-hosted solutions like Ollama or llama.cpp. The platform introduces a concept of “assistants,” which abstract prompting into reusable tools for tasks like translation, summarization, or document analysis, making it easier for non-technical users to leverage AI capabilities. It also incorporates advanced features such as retrieval-augmented generation, plugin extensibility, and support for multiple data sources, allowing users to integrate their own files and knowledge bases into conversations.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 3
    Outlines

    Outlines

    Structured Outputs

    ...Developers can require predefined choices, primitive Python types, Pydantic models, JSON schemas, regular expressions, function signatures, or custom grammars. The same interface works with local models, inference servers, and hosted APIs from several providers. Supported integrations include Transformers, llama.cpp, vLLM, Ollama, OpenAI, and Gemini. It also provides reusable prompt templates and application abstractions for packaging prompts with output constraints. Outlines is designed for extraction, classification, tool calling, automation, and other production workflows that require valid machine-readable results.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Hy-MT2
    Hy-MT2 is a family of fast-thinking multilingual translation models built for complex real-world translation scenarios. It includes 1.8B, 7B, and 30B-A3B model sizes, giving users options for lightweight deployment, stronger general performance, or MoE-based capacity. The models support translation across 33 languages and are designed to follow detailed translation instructions in multiple languages. They can handle tasks involving terminology, style, personalization, delimiters, structured...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Get Started Free
  • 5
    OpenAI Sublime Text Plugin

    OpenAI Sublime Text Plugin

    First class Sublime Text AI assistant with gpt-5, Opus 4.6, Gemini 3

    OpenAI Sublime Text is a full-featured AI assistant plugin designed to bring advanced language model capabilities into the Sublime Text editor with a deeply integrated user experience. It supports a wide range of providers, including OpenAI, Anthropic, Google Gemini, and local backends like Ollama or llama.cpp, making it highly flexible for different deployment scenarios. The plugin offers both chat-based interaction and inline assistance through “phantoms,” which display non-intrusive responses directly within the editor view. It allows users to send selected code, entire files, or diagnostic outputs as context, enabling more precise and relevant responses. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 6
    GPUStack

    GPUStack

    Performance-optimized AI inference on your GPUs

    GPUStack is an open-source GPU cluster management platform designed to simplify the deployment and operation of artificial intelligence models across heterogeneous hardware environments. The system aggregates GPU resources from multiple machines into a unified cluster so developers and administrators can run large language models and other AI workloads efficiently across distributed infrastructure. Instead of requiring complex orchestration systems such as Kubernetes, GPUStack provides a...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    Orpheus TTS

    Orpheus TTS

    Towards Human-Sounding Speech

    Orpheus TTS is a state-of-the-art open-source text-to-speech system built on a Llama-3B backbone, treating speech synthesis as a large language model problem instead of a traditional TTS pipeline. It is designed to produce human-like speech with natural intonation, emotion, and rhythm, targeting quality comparable to or better than many closed-source systems. The project ships both pretrained and finetuned English models, as well as a family of multilingual models released as a research...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 8
    GLM-4

    GLM-4

    GLM-4 series: Open Multilingual Multimodal Chat LMs

    GLM-4 is a family of open models from ZhipuAI that spans base, chat, and reasoning variants at both 32B and 9B scales, with long-context support and practical local-deployment options. The GLM-4-32B-0414 models are trained on ~15T high-quality data (including substantial synthetic reasoning data), then post-trained with preference alignment, rejection sampling, and reinforcement learning to improve instruction following, coding, function calling, and agent-style behaviors. The...
    Downloads: 10 This Week
    Last Update:
    See Project
  • 9
    LoLLMs Hub Fortress

    LoLLMs Hub Fortress

    A proxy server for multiple ollama instances with Key security

    LoLLMs Hub Fortress is a high-performance AI orchestration platform designed to unify multiple large language model backends into a single, secure, and scalable API layer. It acts as a central gateway that connects different inference engines such as Ollama, llama.cpp, vLLM, and OpenAI-compatible services, allowing them to function as interchangeable compute nodes within one system. The architecture is built around a hierarchical “master and slave” hub model, enabling distributed deployments where multiple machines or clusters can be managed through a single entry point. This design allows organizations to scale horizontally, combining local hardware, cloud resources, and specialized inference servers into a unified infrastructure. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Train ML Models With SQL You Already Know Icon
    Train ML Models With SQL You Already Know

    BigQuery automates data prep, analysis, and predictions with built-in AI assistance.

    Build and deploy ML models using familiar SQL. Automate data prep with built-in Gemini. Query 1 TB and store 10 GB free monthly.
    Try Free
  • 10
    ConfiChat

    ConfiChat

    Lightweight, standalone, multi-platform, and privacy focused local LLM

    ...Built as a standalone application without heavy dependencies, it emphasizes simplicity, portability, and ease of setup across multiple platforms including desktop and mobile environments. The tool supports local models such as Ollama and llama.cpp for fully offline operation, while also allowing integration with cloud APIs like OpenAI and Anthropic for access to more advanced capabilities. A key differentiator is its optional encryption of chat history and assets, ensuring that sensitive data can remain secure even when stored locally. Conversations are managed as local JSON files, giving users transparency and direct control over their data. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    FreedomGPT

    FreedomGPT

    React and Electron-based app that executes the FreedomGPT LLM locally

    ...The app's setup is simple, and it includes clear installation guides for both macOS and Windows platforms, as well as detailed instructions for building necessary libraries like llama.cpp.
    Downloads: 7 This Week
    Last Update:
    See Project
  • 12
    llama-cpp-linux-cuda

    llama-cpp-linux-cuda

    https://www.cnblogs.com/Genius-Society/p/20670122

    支持 CUDA 的 linux 版本 llama.cpp 编译成品
    Downloads: 2 This Week
    Last Update:
    See Project
  • 13
    llama2.c

    llama2.c

    Inference Llama 2 in one file of pure C

    ...The goal of llama2.c is to demonstrate how a compact and transparent implementation can perform meaningful inference even with small models, emphasizing simplicity, clarity, and accessibility. The project builds upon lessons from nanoGPT and takes inspiration from llama.cpp, focusing instead on minimalism and educational value over large-scale performance.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 14
    KoboldCpp

    KoboldCpp

    Run GGUF models easily with a UI or API. One File. Zero Install.

    KoboldCpp is an easy-to-use AI text-generation software for GGML and GGUF models, inspired by the original KoboldAI. It's a single self-contained distributable that builds off llama.cpp and adds many additional powerful features.
    Leader badge
    Downloads: 432 This Week
    Last Update:
    See Project
  • 15
    FreeAskInternet

    FreeAskInternet

    FreeAskInternet is a completely free running search aggregator

    ...The backend crawls selected pages, supplies their content to a configured model, and streams a synthesized response to the chat interface. It supports several hosted model services as well as custom local backends such as Ollama and llama.cpp. The web interface is designed for desktop and mobile use and exposes the service only on the local machine by default. Docker Compose provides a straightforward deployment path without requiring a dedicated GPU. The project is still described as early-stage, and actual privacy depends partly on whether users select a local or external model provider.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Amica

    Amica

    Amica is an open source interface for interactive communication

    ...Under the hood, Amica leverages modern web and desktop technologies: three.js and three-vrm for 3D rendering, Transformers.js for running models in the browser, Whisper and Silero VAD for speech recognition and voice-activity detection, and a variety of LLM backends such as llama.cpp servers, ChatGPT-compatible APIs, Ollama, KoboldCpp, and others. It also integrates multiple text-to-speech providers, including ElevenLabs, OpenAI, Coqui, RVC, and AllTalkTTS.
    Downloads: 7 This Week
    Last Update:
    See Project
  • 17
    Alpaca Electron

    Alpaca Electron

    The simplest way to run Alpaca on your own computer

    Alpaca Electron is built from the ground up to be the easiest way to chat with the alpaca AI models. No command line or compiling is needed. Only windows is currently supported for now. The new llama.Cpp binaries that support GGUF have not yet been built for other platforms.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 18
    llama2-webui

    llama2-webui

    Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere

    Running Llama 2 with gradio web UI on GPU or CPU from anywhere (Linux/Windows/Mac).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    TurboPilot

    TurboPilot

    Open source large-language-model based code completion engine

    TurboPilot is a self-hosted copilot clone that uses the library behind llama.cpp to run the 6 Billion Parameter Salesforce Codegen model in 4GiB of RAM. It is heavily based and inspired by on the fauxpilot project. This is a proof of concept right now rather than a stable tool. Autocompletion is quite slow in this version of the project. Feel free to play with it, but your mileage may vary.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    LlamaGPT

    LlamaGPT

    Self-hosted ChatGPT-like chatbot powered by Llama models locally

    ...It supports models such as Llama 2 and Code Llama, allowing users to perform both general conversation and programming-related tasks. It integrates components built around the llama.cpp ecosystem to efficiently run models on consumer hardware. It can be deployed using containerized setups and supports environments ranging from personal computers to self-hosted servers.
    Downloads: 8 This Week
    Last Update:
    See Project
  • 21
    Chinese-LLaMA-Alpaca-2 v2.0

    Chinese-LLaMA-Alpaca-2 v2.0

    Chinese LLaMA & Alpaca large language model + local CPU/GPU training

    This project has open-sourced the Chinese LLaMA model and the Alpaca large model with instruction fine-tuning to further promote the open research of large models in the Chinese NLP community. Based on the original LLaMA , these models expand the Chinese vocabulary and use Chinese data for secondary pre-training, which further improves the basic semantic understanding of Chinese. At the same time, the Chinese Alpaca model further uses Chinese instruction data for fine-tuning, which...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 22
    LlamaChat

    LlamaChat

    Chat with your favourite LLaMA models in a native macOS app

    Chat with your favourite LLaMA models, right on your Mac. LlamaChat is a macOS app that allows you to chat with LLaMA, Alpaca, and GPT4All models all running locally on your Mac.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    LLaMA.go

    LLaMA.go

    llama.go is like llama.cpp in pure Golang

    llama.go is like llama.cpp in pure Golang. The code of the project is based on the legendary ggml.cpp framework of Georgi Gerganov written in C++ with the same attitude to performance and elegance. Both models store FP32 weights, so you'll needs at least 32Gb of RAM (not VRAM or GPU RAM) for LLaMA-7B. Double to 64Gb for LLaMA-13B.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Alpaca.cpp

    Alpaca.cpp

    Locally run an Instruction-Tuned Chat-Style LLM

    ...This combines the LLaMA foundation model with an open reproduction of Stanford Alpaca a fine-tuning of the base model to obey instructions (akin to the RLHF used to train ChatGPT) and a set of modifications to llama.cpp to add a chat interface. Download the zip file corresponding to your operating system from the latest release. The weights are based on the published fine-tunes from alpaca-lora, converted back into a PyTorch checkpoint with a modified script and then quantized with llama.cpp the regular way.
    Downloads: 10 This Week
    Last Update:
    See Project
  • 25
    SuperGemma4

    SuperGemma4

    Fast uncensored Gemma model optimized for local chat and coding

    ...Unlike raw base models, it inherits improvements from the SuperGemma Fast line, resulting in better performance in logic, coding, and real-world text workflows. The model is packaged in GGUF format for efficient use with llama.cpp and has been specifically tested on Apple Silicon hardware, delivering high token speeds and smooth local inference. A neutral chat template is embedded to prevent prompt misrouting issues, ensuring consistent responses without unintended shifts into coding or tool-use modes.
    Downloads: 0 This Week
    Last Update:
    See Project