LLaMA INT8 is an experimental fork of Meta’s original LLaMA code for running large models with reduced memory through 8-bit quantization. It relies heavily on bitsandbytes and the LLM.int8() technique for compressed inference. The implementation quantizes weights on the host and loads them incrementally to avoid large memory spikes. Parallelism constructs from the original code were removed to simplify single-machine execution. The author demonstrated LLaMA 13B running within roughly 24 GiB of GPU memory on an RTX 4090. Repetition-penalty controls and supporting utilities are included for text generation. The project is now mainly historical because its maintainer recommends Hugging Face Transformers for serious modern inference and training workloads.

Features

  • 8-bit LLaMA inference
  • bitsandbytes and LLM.int8() integration
  • Incremental model weight loading
  • Host-side weight quantization
  • Reduced single-GPU memory requirements
  • Configurable repetition penalty

Project Samples

Project Activity

See All Activity >

Categories

AI Models

Follow LLaMA INT8

LLaMA INT8 Web Site

Other Useful Business Software
One Monitoring Tool for IT, OT and Cloud | Free Trial Icon
One Monitoring Tool for IT, OT and Cloud | Free Trial

Vendor-agnostic monitoring across on-prem servers, cloud platforms and OT devices, all in one dashboard. No more tool sprawl.

Modern infrastructure spans data centers, cloud platforms and factory floors, and every blind spot between them is a risk. PRTG supports SNMP, WMI, SSH and other standard protocols to monitor IT, OT and hybrid environments through one customizable dashboard. Build the views your team needs, from network health to application performance, without switching tools. Try PRTG free for 30 days now.
Try PRTG Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of LLaMA INT8!

Additional Project Details

Programming Language

Python

Related Categories

Python AI Models

Registered

2 hours ago