arXiv NLP
techCenter
SGD-KV: Summarization Guided KV Cache Compressiontranslating…
arXiv:2609.03235v1 Announce Type: new
Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics…
Keywords#KV#KV Cache#Cache Compression#Guided#Guided KV
Comments
Sign in to join the discussion.
Loading comments…