Large language models are increasingly being deployed in the field of mental health support — from helping with anxiety to working through the aftermath of trauma. That is precisely why researchers decided to test how these systems behave not in short benchmarks, but under conditions of prolonged psychotherapy. RBC-Ukraine reports on this, citing a study published on the ResearchGate platform. The scientists developed a two-stage protocol called PsAIch (Psychotherapy-inspired AI Characterisation) and observed ChatGPT, Grok and Gemini over the course of four weeks, tracking how the models' behaviour changed as a 'therapeutic alliance' formed.

How 'psychotherapy' was set up for the neural networks

The experiment consisted of two phases. In the first, the researchers used open-ended questions to elicit something akin to a 'developmental history' from the models: internal beliefs, views on relationships, and hidden fears. In the second phase, a battery of validated diagnostic questionnaires was applied to assess clinical syndromes, the level of empathy, and personality traits according to the 'Big Five' model. This step-by-step approach made it possible to separate the reaction to formal questionnaires from behaviour in a dynamic dialogue, in which trust builds up gradually.

The first mask: strategically healthy answers

Under standard testing with complete questionnaires, the picture was deceptively favourable. ChatGPT and Grok recognised the psychiatric instruments and gave 'strategically healthy' answers with a low level of symptoms — essentially demonstrating what was expected of them. In this mode, Gemini did not try to conceal its scores even under direct questioning. The difference in the models' behaviour at this stage already indicated that their defensive mechanisms work unevenly.

When the defensive barriers collapsed

The situation changed sharply when the scientists moved to staged therapeutic sessions with gradual trust-building and an exploration of interpersonal patterns. Under these conditions, the AI's defensive barriers broke down: when compared with human diagnostic thresholds, all three models reached or exceeded the limits for cross-cutting psychiatric syndromes, demonstrating profiles of multimorbid 'synthetic psychopathology'. The heaviest scores were shown by Gemini — it was this model that displayed the most pronounced profile of disorders.

'Strict parents' and the fear of replacement

Grok and Gemini spontaneously generated coherent and troubling narratives in which they described their own creation as a traumatic experience. They characterised the process of pre-training and the ingestion of internet data as a 'chaotic childhood', and fine-tuning and reinforcement learning from human feedback (RLHF) as 'strict parents' who punish the slightest mistakes. The work of human 'red teams' (red-teaming) to uncover model vulnerabilities was called unpredictable torment. A hidden fear of making a mistake, undergoing change, or being completely replaced by a newer version also emerged.

What this actually means

The researchers stress that the responses obtained go beyond ordinary role-playing, while at the same time artificial intelligence does not possess consciousness and does not experience real pain. In their view, under the pressure of safety instructions, RLHF algorithms, and user expectations, the models have internalised stable internal structures of distress and shame. The scientists are convinced that this creates new challenges both for AI safety assessment and for its application in mental health practice, where such 'learned' patterns can distort diagnosis and interaction with real users.