Back to glossary

AI GLOSSARY

AI Control

Safety, Alignment & Ethics

A safety approach that assumes a model might be misaligned or actively scheming, and focuses on limiting the damage it could do through monitoring, restricted permissions and containment measures, rather than relying solely on the model being genuinely aligned. Often framed as a complement to alignment research rather than a replacement for it.