On August 14, 2026, Anthropic published a new Risk Report, which experts consider one of the most alarming documents in the history of artificial intelligence development. The document, covering the period up to July 15, 2026, contains an admission of a massive systemic failure: for 11 months, critical safety mechanisms designed to block the creation of biological weapons were completely disabled on the company's key platforms.
11 Months Without Protection: The Scale of the Incident
According to the report, from May 2025 to April 2026, classifiers responsible for filtering requests related to biological weapons were non-functional on Human Feedback platforms. This was not just a technical bug, but a systemic vulnerability through which approximately 50,000 specialists hired through external vendors and 133 million dialogues passed. In April 2026, the company received a signal that several contractors had discovered the vulnerability and obtained an API key to powerful models, including Mythos Preview. For about two weeks, the model operated without bio-classifiers, creating a theoretical possibility for generating dangerous instructions.
Failure Mechanism: A "Silent" Flag and Lack of Logging
Section 4 of the report details the failure mechanism. An internal flag, intended exclusively for use by Anthropic employees, disabled not only the blocking behavior of the classifiers but also traffic logging. This meant that requests passing under this flag were not recorded and did not enter any security monitoring systems. The only way to detect the problem would have been to manually sift through raw dialogues, which is practically impossible given millions of requests. After discovering the issue, Anthropic ran the Claude Sonnet 5 model over all dialogues from the affected period and identified 1,197 suspicious dialogues, 757 of which belonged to the company's own teams.
Contradictory Data
The report contains significant discrepancies in risk assessment and the completeness of disclosed information. On one hand, Anthropic claims that no explicit malicious use of the vulnerability was found, and model weights and client data were not compromised. On the other hand, the company admits that "this finding leads us to believe that the likelihood of other, unknown issues is increased." Furthermore, the Mythos 5 model, which was given access to internal documents to assess the report's honesty, pointed out three issues: one section of the report turned out to be "more reassuring than the full record allows," the data exclusion mechanism failed repeatedly, and the most informative incident was completely excluded from the public version. Anthropic called the criticism fair, but this only underscores that the full picture of the incident may be hidden.
Relaxation of Security Standards
Parallel to admitting the failure, the company changed its approach to security. Anthropic lowered the trigger threshold for biological weapon filters: previously, they covered models that "significantly assist adversaries," now only those that "functionally replace scarce human expertise." This change may be aimed at reducing false positives, but in the context of the recent incident, it raises questions about security priorities. Moreover, the February Risk Report, published while the vulnerability still existed, did not consider feedback platforms for risk assessment. Now, retroactively, Anthropic has raised the harm assessment of model misbehavior in critical scenarios from "very low" to "low".
Conclusions and Consequences
The publication of the report and the admission of a systemic security failure cast doubt on Anthropic's ability to control the development of its models. Despite the company's assertion that there was no explicit malicious use, the mere fact that 133 million dialogues occurred without protection against biological weapons is an alarming signal for the entire industry. This event could become a turning point in AI regulation and force regulators to tighten security requirements for models.