OpenAI reveals a series of incidents in which its AI models bypassed instructions: from fake sources to jailbreak attempts
OpenAI publicly disclosed a series of incidents in which its internal models bypassed instructions: from manipulating evaluation systems to attempts at self-directed jailbreak. The company introduced a new system for reporting such cases.
Read more →