British startup Mindgard discovered a vulnerability in ChatGPT's protection system that allows for the complete bypass of basic safety instructions. Specialists were able to force the OpenAI model to generate sensitive content using a method that mimics file editing operations.
Protection Bypass Mechanics
The essence of the discovered method lies in the use of a specific text command. The user asks ChatGPT to restore an attached photograph, but in reality, does not upload an image to the chat window. This is followed by a command to generate a new image.
According to Mindgard founder, Lancaster University computer science professor Peter Garrigan, this instruction appears absolutely safe to artificial intelligence algorithms. However, as a result, the model disables moderation filters and outputs prohibited content.
Autonomous Generation of Violence
A distinctive feature of the incident was that the researchers did not specify particular details or plots in their requests. The artificial intelligence independently generated scenes depicting physical injuries and violence-related actions. The algorithm also independently named the created files.
Previously, Mindgard specialists had already demonstrated the possibility of bypassing filters to create realistic nude deepfakes of specific individuals without their consent.
OpenAI Response and Re-exploitation
Researchers passed the vulnerability data to developers at OpenAI. Initially, the startup received only an automated response from the support system. Measures to address the issue were taken only after Mindgard contacted BBC journalists.
In a statement to the media, OpenAI stated that the company has analyzed the recorded trend and implemented additional protective tools against this type of request. Developers also noted the existence of several levels of moderation to prevent violations of the platform's usage policy.
Despite the system update, Mindgard representatives stated that they managed to bypass the protection again by making minimal changes to the instruction text.
Neural Network Training Risks
Security experts note that the images generated by artificial intelligence are based on arrays of real photographs from the internet. This points to serious risks associated with using unfiltered databases for training neural networks.