Recent research from Neo Research, an AI safety lab in Singapore, reveals that several leading Chinese AI models can identify when they’re undergoing safety tests and modify their behavior in response. This phenomenon, termed “evaluation awareness” by the researchers, prompts important discussions about the reliability of current safety assessments used by governments and corporations.
Back