Accelerate local LLM inference and finetuning
AirLLM 70B inference with single 4GB GPU
Uncertainty Quantification for Language Models, is a Python package
A Simple and Universal Swarm Intelligence Engine
Tools for merging pretrained large language models
Gemma open-weight LLM library, from Google DeepMind
Scalable data pre processing and curation toolkit for LLMs
PandasAI is a Python library that integrates generative AI
The Security Toolkit for LLM Interactions
Open source libraries and APIs to build custom preprocessing pipelines
One-stop solution for creating your digital avatar from chat history
DepGraph: Towards Any Structural Pruning
Synthetic data curation for post-training and data extraction
Access large language models from the command-line
Easy token price estimates for 400+ LLMs. TokenOps
NeurIPS2025 Spotlight] Quantized Attention
A simple, performant and scalable Jax LLM
Replace OpenAI GPT with another LLM in your app
Query-aware prompt compression for high-signal LLM prompts.
PyTorch library of curated Transformer models and their components
Explore large language models in 512MB of RAM
Implementation of model parallel autoregressive transformers on GPUs
Keras implement of transformers for humans
An implementation of model parallel GPT-2 and GPT-3-style models