LLaMA INT8 is an experimental fork of Meta’s original LLaMA code for running large models with reduced memory through 8-bit quantization. It relies heavily on bitsandbytes and the LLM.int8() technique for compressed inference. The implementation quantizes weights on the host and loads them incrementally to avoid large memory spikes. Parallelism constructs from the original code were removed to simplify single-machine execution. The author demonstrated LLaMA 13B running within roughly 24 GiB of GPU memory on an RTX 4090. Repetition-penalty controls and supporting utilities are included for text generation. The project is now mainly historical because its maintainer recommends Hugging Face Transformers for serious modern inference and training workloads.

Features

  • 8-bit LLaMA inference
  • bitsandbytes and LLM.int8() integration
  • Incremental model weight loading
  • Host-side weight quantization
  • Reduced single-GPU memory requirements
  • Configurable repetition penalty

Project Samples

Project Activity

See All Activity >

Categories

AI Models

Follow LLaMA INT8

LLaMA INT8 Web Site

Other Useful Business Software
Veeam Data Platform v13.1 - Get Your Free Trial Icon
Veeam Data Platform v13.1 - Get Your Free Trial

Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
Try it Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of LLaMA INT8!

Additional Project Details

Programming Language

Python

Related Categories

Python AI Models

Registered

2 hours ago