arXiv CS AI
techCenter
HyQuant: Hybrid-Precision Quantization for LLM Attentiontranslating…
arXiv:2608.27875v1 Announce Type: new
Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing…
Keywords#Quantization#Attention#LLM#Hybrid-Precision#Hybrid-Precision Quantization