Cisco research tested 15 frontier AI models from OpenAI, Anthropic, Google, and xAI, finding multi-turn attacks achieved safety bypass rates as high as 88%.
An Anthropic cofounder traveled to the Vatican and told Pope Leo XIV that researchers are finding "unsettling" things inside AI models, a rare ethics signal.