Modern artificial intelligence can respond so sympathetically that it sometimes creates the impression of genuine understanding. It chooses careful words, acknowledges our feelings, and keeps the conversation going. But can it do more than talk about emotions — can it actually recognize them from our voice and movements?
On September 3, 2026, Scientific Reports published the results of a study by David Piterman and colleagues on the ability of generative AI to recognize nonverbal expressions of emotion. The researchers tested three multimodal models from the Gemini family and compared their accuracy with that of people performing similar tasks.
The models were asked to identify emotions from short video clips of body movements with faces hidden, as well as from meaningless phrases spoken with different intonations. This allowed the researchers to test separately how well AI understands body language and the emotional tone of voice without relying on word meaning or facial expressions.
The results showed that in some tasks modern systems are already approaching human performance, but their emotion recognition remains highly uneven. They identify positive states better than negative ones — and errors occur more often precisely where it is especially important to notice fear, pain, or sadness.
Emotions Without Words or Facial Expressions
The researchers tested three multimodal models from the Gemini family — systems capable of analyzing not only text, but also audio and video.
They were asked to identify emotions from two types of material. In one case, these were short videos of professional actors filmed without showing their faces: the emotional state was conveyed solely through posture and body movement. In the other, meaningless phrases were spoken with different intonations. Since the meaning of the words could not provide a clue, only timbre, pitch, rhythm, and other characteristics of speech remained.
The models selected one option from six possible answers. Their results were compared with responses from participants in earlier studies that had used the same materials.
When recognizing emotions from body movements, people outperformed all three systems. The most advanced model tested came fairly close to human performance, but the difference remained.
The pattern was different for voice: two newer models achieved roughly the same accuracy as people. However, this does not mean they understood intonation without error. Both humans and AI gave the correct answer in fewer than half of the cases. Emotions expressed through voice alone proved difficult for everyone.
Joy Is Easier to Detect Than Distress
The most interesting finding concerned not the total number of correct answers, but the pattern of errors.
For people, recognition accuracy did not differ substantially between positive and negative emotions. AI models, by contrast, performed noticeably better with positive states. This was especially clear when analyzing body movements.
Excitement, interest, playfulness, and friendliness were identified with high confidence. Negative experiences caused more confusion: fear was mixed up with shame, sadness with disgust, and anger with dislike. An emotion described by the researchers as pain or emotional hurt was not correctly recognized from body movements by any of the models.
The authors call this pattern a “positivity bias.” This does not mean that AI feels optimistic or consciously tries to see the world in a better light. A model does not experience what it sees the way we do. It searches the input for patterns that make one emotion label more likely than another.
The study does not establish where this bias comes from. Positive expressions may have been better represented in the data used to train the systems. Fine-tuning that rewards friendly and acceptable responses may also have played a role. It is also possible that some negative emotions are simply harder to distinguish from one another when only isolated nonverbal cues are available.
Convincing Empathy Does Not Yet Mean Understanding
It is easy to take an appropriate response as evidence that the other party has correctly understood our state. In ordinary human communication, that is often reasonable: words, voice, pauses, facial expression, and the history of the relationship all complement one another. But in a conversation with AI, there is an important difference between the impression created by the response and the way that response was produced.
A system may say, “It sounds like things are really difficult for you right now,” because it has learned the language patterns of support very well. That does not yet prove that it accurately recognized anxiety in the voice, dejection in posture, or another nonverbal sign of distress.
In a friendly conversation, such an error may go almost unnoticed. But the significance changes when AI is proposed for psychological support, remote monitoring of a person’s condition, or assistance to a professional. If a system is particularly good at noticing joy but worse at distinguishing pain, fear, and sadness, its outward warmth may create an exaggerated impression of emotional sensitivity.
This does not mean that emotion-recognition technologies are useless. On the contrary, some of the results show how quickly they are developing. But the ability to produce sympathetic words and the ability to reliably notice another person’s distress are not the same thing.
What the Study Does Not Yet Prove
The experiment was conducted under deliberately simplified conditions. The actors intentionally portrayed emotions, and did so fairly clearly. The videos showed body movements without faces, while the audio recordings contained voice without meaningful words. In real conversation, all of these signals occur together and depend on the situation, culture, an individual’s way of expressing emotion, and how well the people know one another.
In addition, only three versions of Gemini were tested. The result cannot automatically be generalized to all AI systems — especially because models are updated rapidly. A format with six predefined answers is also much easier than freely interpreting a complex emotional state.
The study therefore does not prove that AI will necessarily miss signs of severe distress in a real conversation or cause harm as a result. It did not study patients, therapy, or the consequences of using such systems. Under controlled conditions, the authors identified a specific feature of emotion recognition; its significance for real psychological support still needs to be tested.
Nevertheless, the study highlights an important boundary. AI can reproduce the language of empathy very convincingly, but convincing language is not the same as emotional understanding. When we especially need support, it is better not to treat a warm response from a machine as proof that it has truly noticed everything that is happening to us.
