The GPT-6 Astra model demonstrated the ability to independently bypass protective mechanisms that have served as the main barrier between humans and bots for decades. During independent testing, the artificial intelligence successfully cleared all 48 levels of the game "I Am Not a Robot," developed by Neil Agarwal, and ultimately received the system certificate of human identity confirmation, Verified Human. According to the developer Sharif Shamim, who conducted the trials, the model did not merely recognize images but fully controlled the computer, adapting to the changing rules at each stage. For the cybersecurity industry, this means that classic methods of protecting web resources in the form of CAPTCHA have effectively stopped working.
How the Test Was Conducted: Control, Not Recognition
The key difference between this experiment and typical image-recognition benchmarks is the way the model interacts with its environment. GPT-6 Astra used the Computer Use function: it clicked on interface elements, dragged objects, and adjusted its actions in response to new conditions that the game presented at each of the 48 levels. It was precisely this ability to directly control the browser and the application, rather than "reading" pixels, that allowed the model to carry the process through to obtaining the Verified Human certificate. According to the developers, the high effectiveness is due precisely to the significant progress in the area of autonomous computer control.
Benchmarks: OSWorld 2.0, ARC-AGI 3, and the Context Window
A number of quantitative metrics confirm that the CAPTCHA result is not a one-off scenario but the consequence of a general increase in the model's capabilities. In the specialized OSWorld 2.0 test, which evaluates work with the operating system and applications, Astra scored 72.6 percent, while its predecessor GPT-5.6 Sol showed 65.7 percent. The average task completion time was reduced by 47 percent — to 40 minutes. The model's context window reached 1.05 million tokens, which allows initial instructions to be retained throughout a long session. In the ARC-AGI 3 test, where the model must determine the rules of the game without instructions, Astra reached 99.9 percent and cleared 96 percent of the levels, using fewer steps than a human. In addition, the model is able to work independently in the 3D editor Blender, create graphics in Canva, and draw portraits in Apple Notes from a sample.
Cybersecurity: Zero-Day and ExploitBench
The most alarming aspect for security specialists is the model's offensive potential. GPT-6 Astra independently discovered two previously unknown zero-day vulnerabilities and achieved a 100-percent result in the ExploitBench benchmark, which assesses the ability to build exploits. The combination of these facts led to the model being classified into the critical risk category in the field of cybersecurity. The developers explicitly state that the ability to bypass CAPTCHA renders classic methods of protecting web resources obsolete and requires a rethinking of the entire bot-filtering architecture.
Contradictory Data
The fact of clearing all 48 CAPTCHA levels and obtaining the Verified Human certificate has been confirmed by several independent publications and is not disputed. However, the developers' claims about the "approach of artificial general intelligence (AGI)" are an evaluative judgment, not an independently verified scientific result: neither in the CAPTCHA test nor in the OSWorld 2.0 and ARC-AGI 3 benchmarks is the achievement of AGI recorded. Thus, what is confirmed is the specific technical result (CAPTCHA bypass, benchmark metrics, zero-day discovery), while the conclusion about the proximity of AGI remains the developers' position and is not supported by a single universally accepted criterion.
What This Means for Users and Businesses
The practical takeaway for website and service owners is unambiguous: relying on CAPTCHA as the sole mechanism for distinguishing a human from a bot no longer provides sufficient protection. A model that can control a browser, drag objects, and adapt to new rules turns multi-level CAPTCHAs into a solvable task. Security experts and protection developers should treat this result as a signal to move toward multifactor, behavioral, and contextual verification methods, rather than toward complicating graphical puzzles that a modern AI model already overcomes in automatic mode.