arXiv NLP
techCenter
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoningtranslating…
arXiv:2609.03430v1 Announce Type: new
Abstract: Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score…
Keywords#KV Cache#Reasoning#Attention#Cache Eviction#Efficient
Comments
Sign in to join the discussion.
Loading comments…