AutoAWQ
JSON →AutoAWQ implements the AWQ (Activation-aware Weight Quantization) algorithm for 4-bit quantization of large language models, achieving up to 2x speedup during inference. The library is now deprecated as of v0.2.9 (April 2025), with vLLM having adopted the technology. Last tested with Torch 2.6.0 and Transformers 4.51.3.
Traffic · last 30 days stale · no recent hits
total hits 19
actors 6 distinct systems
last hit 16d ago AhrefsBot
top countries 🇺🇸 United States · 🇸🇬 Singapore · 🇩🇪 Germany · 🇬🇧 United Kingdom · 🇨🇦 Canada
API endpoints
full doc /v1/registry/autoawq
install /v1/registry/autoawq/install
compatibility /v1/registry/autoawq/compatibility