arXiv Machine Learning
techCenter
A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Modelstranslating…
1 min readUnknownarXiv Digital Media
arXiv:2608.26926v1 Announce Type: new
Abstract: Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models…