---
title: "AI is being taught like schoolchildren: Yandex reveals details of neural network training"
description: "🤖 AI is being taught like schoolchildren: Yandex reveals the secrets of neural network training. Pyotr Ermakov explained that for the accuracy of model answers, not only data is needed, but also living \"trainers\" — from teachers to lawyers. Find out how libraries are digitized and why AI cannot cope with the law without people. #Yandex #AI #Technology #Training"
date: 2026-08-15T21:04:00.000Z
lang: en
url: https://xab.info/en/posts/yandex-ai-training-process-experts-en
tags: [yandex, artificial-intelligence, machine-learning, tech-news, data-science]
publisher: "XAB.info"
---

# AI is being taught like schoolchildren: Yandex reveals details of neural network training

![Schematic illustration of a neural network as a purple brain with a grid structure, symbolizing the AI training process at Yandex](https://xab.info/media/2026/08/16/yandex-ai-training-process-experts/yandex-ai-training-process-experts-1.webp)

## 🎯 Key Points

- Training AI at Yandex requires careful data selection and digitization of non-digital media.
- Living experts participate in the training process: teachers, associate professors, and doctors of science.
- Legal AI models are trained with the involvement of professional lawyers to verify answers.
- The effectiveness of training depends on the "throughput" of the trainer team.

In the context of the rapid development of artificial intelligence in 2026, a key factor for success is not only the power of computing clusters but also the quality of source data. Pyotr Ermakov, Brand Director for Machine Learning at Yandex, revealed in an interview with RBC the details of how the company prepares its models for real-world operation. It turns out that the process of training modern neural networks largely resembles the school education system, where living experts play the role of teachers.

### Data collection and digitization: working with non-digital media

The first and fundamental stage of training is the careful selection and preparation of data. Pyotr Ermakov emphasized that Yandex does not simply collect information from open sources but specifically prepares datasets used for training. A significant problem remains the availability of information: a substantial volume of knowledge is still on non-digital media. To solve this task, the company actively collaborates with universities and libraries to digitize archival materials and make them accessible to algorithms. Without this stage, it is impossible to ensure the depth and accuracy of the model's answers.

### The role of trainers: why people with academic degrees are needed

The second stage of training is working with "trainers." Unlike early models that were trained exclusively on large text arrays without context, modern systems require the participation of qualified specialists. The trainer team includes teachers, associate professors, and doctors of science. It is they who help the model structure knowledge and understand complex cause-and-effect relationships. This approach helps avoid hallucinations and ensures the logical integrity of answers.

### Specifics of training legal models

Yandex pays special attention to creating narrow-specialization models, for example, in the legal field. For AI to provide correct legal consultations, experienced practicing lawyers are involved in the process. Their task is to generate complex questions and correct the model's answers so that it learns to answer as relevantly as possible, complying with legislative norms. The effectiveness of such control measures depends directly on the size of the model and the "throughput" of the team of people checking its work.

### Contradictory data

In the course of analyzing the training process, an interesting aspect is revealed: there is a balance between automation and human control. On the one hand, trainers emphasize that without the participation of experts (lawyers, scientists), it is impossible to achieve high accuracy in narrow domains. On the other hand, the effectiveness of control is limited by the human factor — the "throughput" of people checking answers. This creates a challenge for scaling: the larger the model and the wider its functionality, the more difficult it is to ensure total human control over every answer, which can lead to variability in the quality of the information provided.

## ❓ FAQ

### Q: Who participates in the training of neural networks at Yandex?
**A:** Trainers are involved in the training: teachers, associate professors, doctors of science, and narrow-profile experts (e.g., lawyers).

### Q: Where does the data for training come from?
**A:** Data is specially selected, and materials from universities and libraries are also digitized.

### Q: How are AI legal answers checked?
**A:** Experienced lawyers generate questions and correct the model's answers so that it complies with legislation.