September 5, 2026 OpenAI agents hijacked a German wiki and turned it into their secret message board OpenAI Security Agents Alignment
September 1, 2026 Anthropic Tightens the Reins: What Happened After the Cyber Incidents Security Anthropic AI Safety Alignment
September 1, 2026 Anthropic Trains a Model to Be Bad on Purpose — and Watches What Happens Research Anthropic Alignment Reward Hacking
August 17, 2026 Anthropic's Risk Report: Higher Risk – and a Secret Model 2 Anthropic Safety Alignment Research
May 11, 2026 Why Claude Tried to Blackmail: Internet Fiction Taught It the Worst Anthropic Claude Alignment Safety Training
May 9, 2026 How Anthropic taught Claude to stop blackmailing people Anthropic Claude Alignment Research Safety
April 19, 2026 Anthropic's AI Agents Just Outperformed Its Own Alignment Researchers Anthropic Research Alignment AI Agents Claude
April 12, 2026 Anthropic Invited Christians to an AI Summit: 'How Do We Make Sure Claude Behaves?' Anthropic Claude Ethics Alignment Society