Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
arxiv.org·6d
📊Quantization
Preview
Report Post

View PDF

Abstract:As large language models have grown larger, low-precision numerical formats such as NVFP4 have become increasingly popular due to the speed and memory benefits they provide. However, to accelerate computation with NVFP4, all matrix multiplication operands–weights and activations in the forward pass, and weights, activations, and gradients in the backward pass–must be quantized to NVFP4, often leading to divergence during training and performance degradation during inference. NVFP4 by evaluating multiple potential scale factors for each block of values. To address this issue, in this work we introduce Four Over Six (4/6), a modification to the NVFP4 quantization algorithm that evaluates two potential scale factors for each bloc…

Similar Posts

Loading similar posts...