MR-GPTQ: Better FP4 Microscaling for LLMs
Автор: AI Research Roundup
Загружено: 2025-10-06
Просмотров: 162
Описание:
In this AI Research Roundup episode, Alex discusses the paper:
'Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization'
This work evaluates MXFP4 and NVFP4 FP4 formats for post-training quantization of LLMs, explaining why standard PTQ struggles due to small group sizes, power-of-two scales, and non-uniform grids. The authors analyze quantization error and show Hadamard rotations can help MXFP4 but hurt NVFP4 with RTN, and that MXFP4’s E8M0 scales induce large errors. They introduce Micro-Rotated-GPTQ (MR-GPTQ) with fused Hadamard rotations, MSE-optimized scale/grid search, activation reordering, and fast GPU kernels (QuTLASS) for Blackwell. Results across Llama-3 and Qwen-3 show FP4 is lossy, but MR-GPTQ narrows the accuracy-performance gap while enabling efficient inference.
Paper URL: https://arxiv.org/abs/2509.23202
#AI #MachineLearning #DeepLearning #LLM #Quantization #FP4 #GPTQ #GPU
Resources:
GitHub: https://github.com/IST-DASLab/FP-Quant
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: