arXiv CS AI
techCenter
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrailstranslating…
arXiv:2608.27580v1 Announce Type: new
Abstract: Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that…
Keywords#Safety Guardrails#LongGuard#Analysis#Failure#Long-Context