Back to glossary
AI GLOSSARY
Constitutional Classifiers
Safety, Alignment & Ethics
A safeguard developed by Anthropic that uses separate classifier models, trained on a written set of principles, to detect and block harmful inputs or outputs in specific risk categories. It extends the idea behind Constitutional AI into a dedicated filtering layer around a deployed model.