LLaMA INT8 is an experimental fork of Meta’s original LLaMA code for running large models with reduced memory through 8-bit quantization. It relies heavily on bitsandbytes and the LLM.int8() technique for compressed inference. The implementation quantizes weights on the host and loads them incrementally to avoid large memory spikes. Parallelism constructs from the original code were removed to simplify single-machine execution. The author demonstrated LLaMA 13B running within roughly 24 GiB of GPU memory on an RTX 4090. Repetition-penalty controls and supporting utilities are included for text generation. The project is now mainly historical because its maintainer recommends Hugging Face Transformers for serious modern inference and training workloads.
Features
- 8-bit LLaMA inference
- bitsandbytes and LLM.int8() integration
- Incremental model weight loading
- Host-side weight quantization
- Reduced single-GPU memory requirements
- Configurable repetition penalty