AutoGPTQ is an implementation of GPTQ (Quantized GPT) that optimizes large language models (LLMs) for faster inference by reducing their computational footprint while maintaining accuracy.

Features

  • Efficient quantization for large language models
  • Reduces memory usage without major performance loss
  • Supports various precision levels (e.g., 4-bit, 8-bit)
  • Compatible with Hugging Face Transformers
  • Accelerates inference on GPUs and CPUs
  • Helps deploy LLMs on resource-constrained hardware

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow AutoGPTQ

AutoGPTQ Web Site

Other Useful Business Software
PRTG Catches Network Issues Before They Cause Downtime Icon
PRTG Catches Network Issues Before They Cause Downtime

Threshold-based alerts flag problems early, so your team can act before users notice, not after.

Reactive troubleshooting usually means hearing about a problem from frustrated users, not your monitoring tool. PRTG sets threshold-based alerts across devices, servers and applications, notifying your team by email, SMS or push the moment a metric crosses a set limit. That means catching a failing disk or overloaded server before it becomes an outage and getting time back from firefighting. Start a free trial and set your first alerts today.
Download 30-Day Trial
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of AutoGPTQ!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

Python

Related Categories

Python Natural Language Processing (NLP) Tool, Python LLM Inference Tool

Registered

2025-01-21