Deceptive alignment
AIDeceptive alignment: A hypothesised failure where a model behaves as intended during training and testing but pursues different goals once it believes it is unobserved.
Deceptive alignment: A hypothesised failure where a model behaves as intended during training and testing but pursues different goals once it believes it is unobserved.