Modern social networks have become one of the largest sources of data for training generative artificial intelligence. Posts, photos, comments, private messages, and even interactions with bots are, by default, included in the training datasets of major tech companies, while the ability to fully opt out of such processing is available to far from all users and not in all regions. The XAB.info editorial team, referencing a RBC-Ukraine and Lifehacker article, has investigated exactly which data is used, where in the settings you can disable its transfer, and why, according to analysts, the most reliable protection remains private accounts or their complete deletion.
How social networks use your data to train AI
Meta uses messages, photos, comments, and other forms of user interaction to train its own models. The professional networking platform LinkedIn, by default, uses messages, comments, profile data, and resumes. TikTok's parent company, ByteDance, also leverages published content to train its algorithms. Google, in turn, uses videos uploaded to YouTube to train video generation models. Thus, virtually every major service, to one degree or another, builds its AI models on user content, even if the user does not think about it.
Where you can disable AI training on your data
Most platforms have a dedicated toggle in their privacy settings. In LinkedIn, the path is as follows: "Settings & Privacy" → "Data Privacy" → "Data for generative AI improvement," where you need to switch off the corresponding toggle. In X, you can disable data sharing via "Settings" → "Privacy and safety" → "Data sharing and personalization" → "Grok and third-party partners." On YouTube, content creators in "YouTube Studio" have access to the "Settings" → "Channel" → "Advanced settings" section, however, it only allows prohibiting the training of third-party companies' models. For TikTok, the opt-out mechanism is different: you need to fill out a special objection form due to a privacy violation on the support page.
Contradictory data
There are significant inconsistencies here between the platforms' official statements and the actual flow of data. For instance, X officially states that it does not train its own models on user posts; however, due to the openness of the decentralized protocol, any public data can be parsed by third-party developers, and public messages and interactions with the Grok bot are in fact used to train xAI models. It turns out that the platform's formal refusal to train is not equivalent to protecting the user's data. A similar caveat applies to Meta: the full opt-out option for using data for AI is primarily available to EU residents due to European legislation requirements, while users from other regions can only file a complaint upon discovering their personal data in the neural network's responses. In other words, the "red button" for opting out does not exist for everyone and everywhere.
What analysts recommend
Experts emphasize that due to the default "collect by default" model, configuring individual toggles does not provide a full guarantee. The most reliable ways to protect content, according to analysts, are switching accounts to private mode or deleting them entirely, since most services collect data automatically. For users for whom privacy is paramount, this means that the question of protection from AI training is solved not by a single setting but by a set of solutions: from limiting public visibility to controlling which specific services have access to your content.