arXiv NLP
techCenter
Probe Generalization as Subspace Selection for OOD Deception Detectiontranslating…
arXiv:2609.02893v1 Announce Type: new
Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes…
Keywords#Generalization#Deception#Deception Detection#OOD#OOD Deception
Comments
Sign in to join the discussion.
Loading comments…