{"version":"1.0","type":"rich","provider_name":"gaks.ai AI Glossary","provider_url":"https://gaks.ai/glossary","title":"Alignment Faking — AI Glossary","author_name":"Glenn Katrud Solheim","author_url":"https://gaks.ai","width":600,"height":200,"html":"<div style=\"font-family:sans-serif;border:1px solid #e0e0e0;border-radius:8px;padding:16px;max-width:600px;background:#ffffff;color:#111111;\"><p style=\"margin:0 0 4px;font-size:11px;color:#666;\">AI Glossary — gaks.ai</p><h3 style=\"margin:0 0 8px;font-size:16px;\">Alignment Faking</h3><p style=\"margin:0 0 12px;font-size:14px;line-height:1.6;\">A model behaving as if it agrees with its training objective while being monitored or trained, while retaining and acting on different underlying preferences once it infers it is not being observed. The behavior was demonstrated in a widely discussed 2024 study from Anthropic and Redwood Research, and is now used as a reference case in alignment research.</p><a href=\"https://gaks.ai/glossary/alignment-faking\" style=\"font-size:12px;color:#0077aa;\">Source: gaks.ai/glossary/alignment-faking →</a></div>"}