UniSVQ: 2-bit Unified Scalar-Vector Quantization (opens in new tab) 🚀ML Inference Content type: Academic

arxiv.org··Covered by ai-brief.liziran.com·Open original

Post-training quantization at the 2-bit level enables low-cost deployment and inference acceleration for large language models (LLMs). Scalar quantization (SQ) and vector quantization (VQ) are two primary quantization methods, however, the former suffers from significant performance degradation, and the latter incurs computational and storage overhead. We propose UniSVQ, a unified 2-bit quantization framework that bridges scalar and vector quant...

Read the original article

Sign in to keep reading the full article.

Sign Up Log In

Cited by 1 article

In other languages

一条证据压成1个token，生成省3-10倍

ai-brief.liziran.com·