Tagged "gptq"
3 articles tagged gptq, 7 August 2026 to 19 September 2026. Newest first.
-
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained
A comprehensive comparison of the major quantization formats used in local LLM deployment, covering GGUF, GPTQ, AWQ, and EXL2 formats and their tradeoffs for on-device inference.
-
Benchmarking Qwen3.8 27B Quantizations: 4-bit Holds Up, 1-bit Collapses
Detailed quantization benchmarks for Qwen3.8 27B reveal that 4-bit quantization maintains strong performance while extreme 1-bit quantization severely degrades output quality, providing practical guidance for practitioners choosing compression levels.
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
Daniel Han explores how aggressive model compression can maintain capabilities, challenging assumptions about size-to-performance tradeoffs in quantization and pruning for local inference.