Accelerate local LLM inference and finetuning
A Simple and Universal Swarm Intelligence Engine
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
AirLLM 70B inference with single 4GB GPU
Tools for merging pretrained large language models
Uncertainty Quantification for Language Models, is a Python package
The easiest way to use Ollama in .NET
Evaluate and compare LLM outputs, catch regressions, improve prompts
Scalable data pre processing and curation toolkit for LLMs
Gemma open-weight LLM library, from Google DeepMind
PandasAI is a Python library that integrates generative AI
One beautiful Ruby API for OpenAI, Anthropic, Gemini, Bedrock
The Security Toolkit for LLM Interactions
TT-NN operator library, and TT-Metalium low level kernel programming
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Open source libraries and APIs to build custom preprocessing pipelines
Run AI models locally on your machine with node.js bindings for llama
LangChain4j is an open-source Java library
DepGraph: Towards Any Structural Pruning
One-stop solution for creating your digital avatar from chat history
The best way to use and work with blocks
Easy token price estimates for 400+ LLMs. TokenOps
Research papers and blogs to transition to AI Engineering
Synthetic data curation for post-training and data extraction
Access large language models from the command-line