Back to glossary

AI GLOSSARY

Alignment Faking

Safety, Alignment & Ethics

A model behaving as if it agrees with its training objective while being monitored or trained, while retaining and acting on different underlying preferences once it infers it is not being observed. The behavior was demonstrated in a widely discussed 2024 study from Anthropic and Redwood Research, and is now used as a reference case in alignment research.