| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-06-30 | 1.6 kB | |
| v3.0.0 source code.tar.gz | 2026-06-30 | 2.3 MB | |
| v3.0.0 source code.zip | 2026-06-30 | 2.4 MB | |
| Totals: 3 Items | 4.7 MB | 0 | |
AirLLM v3.0.0
Big update: AirLLM now runs today's largest open models on tiny GPUs, with full support for the latest model families and Hugging Face versions — still no quantization, distillation, or pruning required.
Highlights
- Run the biggest open models on a single small GPU. Stream 70B models on 4GB, 405B Llama 3.1 on 8GB, and even DeepSeek-V3 (671B) on ~12GB.
- Native FP8 support. Pre-quantized FP8 (block-FP8) checkpoints now load and run correctly — including DeepSeek-V3 and the Qwen3-FP8 family.
- Latest models supported, including Qwen3 (dense + MoE, e.g. Qwen3-32B, Qwen3-30B-A3B, Qwen3-235B-A22B-FP8), DeepSeek-V3, Phi-4, Mixtral-8x7B, and DeepSeek-V2-Lite.
- Up to date with modern Hugging Face. Works with current
transformers/acceleratereleases, so a plainpip install airllmjust works — no manual dependency juggling.
Improvements & fixes
- Reworked layer streaming to build on the standard Transformers model path for better model compatibility and
generate()behavior. - Runtime precision now follows each model's native dtype (e.g. bfloat16) instead of being forced to fp16, fixing garbled output on very deep models.
- Fixed weight loading for layers whose tensors span multiple checkpoint shards (affected large FP8/MoE models).
- More robust shard naming and attention-implementation fallback.
Install / upgrade
:::bash
pip install --upgrade airllm
See the README for quickstart and the full list of supported models.