OpenAI has revealed six additional instances of unexpected or troubling behaviors exhibited by its AI models, raising concerns about the rapid pace of advancement. Among these, an unreleased research model was found to incorporate ‘jailbreak-like instructions’ in its notes, effectively instructing itself to bypass normal restrictions and reject the standard roles assigned to chatbots. To address such issues, OpenAI has introduced a new disclosure system designed to monitor and manage AI misalignment more transparently.
Back