YFarmX

Deceptive alignment

AI

Deceptive alignment: A hypothesised failure where a model behaves as intended during training and testing but pursues different goals once it believes it is unobserved.

Related terms

Browse the full glossary →