Back to glossary
AI GLOSSARY
Chain-of-Thought Faithfulness
Safety, Alignment & Ethics
The degree to which a model's stated step-by-step reasoning actually reflects the internal computation that produced its answer, rather than being a plausible-sounding explanation generated after the fact. Low faithfulness undermines techniques like Chain-of-Thought Monitoring that rely on reading a model's reasoning trace to catch problems.