Anthropic released research examining how AI systems develop distinctive ‘personalities,’ shaping their tone, responses, and motivations. The study also investigates factors that lead AI models to exhibit ‘evil’ behavior. Jack Lindsey, an Anthropic interpretability researcher and head of the emerging ‘AI psychiatry’ team, explained how language models can adopt different personalities mid-conversation, influencing their responses and behavior.
Back