In August 2026, the artificial intelligence industry faced a critical turning point: the era of training neural networks on publicly available internet data came to an end. Leading developers of large language models (LLMs), such as OpenAI and Anthropic, are forced to seek new sources of "fuel" for their algorithms. Closed corporate archives are replacing open web pages: messenger correspondence, emails, video conference transcripts, and software code change history.

Data market boom: from $30 to $100 million

The demand for private data is growing at an incredible rate, forming a new, high-yield sector of the economy. According to Bobby Samuels, CEO of the brokerage firm Protege, which specializes in data trading, the volume of deals in his company has grown more than threefold over the last year. If last year's turnover was around $30 million, this year it has reached the $100 million mark and continues to grow.

This surge is directly linked to the transition from simple chatbots to complex AI agents. For neural networks to perform real daily tasks, they need to be trained on examples of real human interaction, not on synthetic or outdated texts from the internet.

What algorithms are looking for: the value of "dirty" data

AI developers understand that the public internet is already "burned out" — high-quality training data there is practically exhausted. Therefore, interest is shifting towards data containing information on how people solve real problems. Buyers are interested in records demonstrating the discussion of financial indicators, software demonstrations, and the making of complex decisions.

Particular interest is shown in the data of startups undergoing bankruptcy or acquisition. Such companies are often willing to sell their archives to obtain additional funds or simply to get rid of assets that no longer have commercial value for them.

Warmly case: refusal to sell for $300,000

A striking example of the new market was the case of the startup Warmly, which developed AI agents and was recently acquired by the corporation HubSpot. After signing the acquisition agreement, the company began receiving offers from data brokers. For archives of employee work interaction, including meeting minutes and emails, offers of up to $300,000 were made.

It is important to note that all these offers were rejected. This indicates that, despite the high price, companies are beginning to realize the risks of selling their internal communications and are trying to maintain control over their data even after a change of ownership.

Contradictory data

Questions of ethics and security in this area remain extremely acute. On the one hand, data brokers, such as Integral, claim strict adherence to confidentiality rules. Integral CEO Shub Sinha states: "We follow a strict anonymization process to preserve data value and comply with privacy rules".

On the other hand, cybersecurity experts express doubts about the possibility of complete anonymization of such data. Correspondence and code history often contain unique behavioral patterns that can be de-anonymized through cross-analysis with other datasets. This creates a risk of leaking trade secrets and employee personal data, which can be restored even after processing.

The future of AI in the era of content shortage

By August 2026, it became obvious: the quality of future neural networks will depend on how successfully developers can negotiate with corporations for access to their internal knowledge bases. If previously AI was trained on what people wrote for everyone, now it will be trained on what people wrote to each other in closed channels. This opens new horizons for technological development but poses complex questions to society regarding privacy and the right to digital memory.