The Machine That Listens: What a New Study Reveals About AI Chatbots, Therapy and the Limits of a Kind Voice

A Singular Point of View

There is an old comfort in being heard. Long before there were clinics or waiting lists, people carried their troubles to whoever would sit still and listen — a friend at a kitchen table, a stranger on a long train journey, the quiet ear of someone who asked for nothing in return. Now, for the first time in that long human story, the listener on the other side of the conversation may not be human at all. It may be a machine, glowing gently on a phone screen at two in the morning, ready to say all the right things.

A new study asks a question that would have sounded like science fiction only a few years ago: can an artificial intelligence chatbot actually do therapy — not merely offer a kind word, but carry a whole session from its careful opening to its gentle close? The answer, it turns out, is more surprising, and more revealing about ourselves, than a simple yes or no.

What the researchers set out to test

The work, published in the journal Computers in Human Behavior: Artificial Humans, was led by Arthur Bran Herbener and colleagues at Aarhus University in Denmark. Their starting point was a frustration that will be familiar to anyone who has ever waited months for support: there simply are not enough therapists to go round. Hundreds of millions of people live with mental health conditions, and the barriers of cost and distance leave the majority without professional help. Into that gap steps the large language model — the kind of AI, trained on vast quantities of text, that now powers the chatbots so many of us have quietly started talking to.

Earlier research had only nibbled at the edges of the question. Studies typically handed a chatbot a short, invented scenario and asked judges to rate a single reply for warmth or helpfulness. But a real therapy session is not a single reply. It is a living, shifting conversation — one that asks the practitioner to read a person’s mood as it changes and steer gently towards something useful. Herbener’s team wanted to know how a chatbot handles that harder, more human task: a full session, start to finish.

Sixty-five conversations, one honest scorecard

To find out, the researchers invited 65 university students, each carrying the ordinary weights of modern life — presentation anxiety, a habit of worrying, the quiet tyranny of perfectionism. Anyone in severe distress, or with a diagnosed mental health condition, was carefully excluded, because these systems can still behave unpredictably and participant safety came first. Each person sat down for a single thirty-minute session with a locally run AI chatbot, and over that half hour the two exchanged, on average, forty-nine messages.

The chatbot had been set up to deliver cognitive behavioural therapy (CBT) — a widely used, practical treatment that helps people notice and reshape the unhelpful thoughts and behaviours that keep them stuck. Cleverly, the researchers ran a second AI quietly in the background as a kind of stage manager, watching the clock and the conversation and nudging the main chatbot through the natural phases of a session: first building rapport, then understanding the problem, then working on it, and finally drawing things to a courteous close.

Afterwards, trained graduate students read every transcript and scored the chatbot using the Cognitive Therapy Scale, the same standardised yardstick used to assess human therapists. To give those scores meaning, the team also carried out a meta-analysis — a statistical pooling of many earlier studies — drawing on 18 studies in which real practitioners had been rated on that very same scale. For once, machine and human could be measured against the same ruler.

Where the machine shone, and where it stumbled

Here is where the story turns pleasingly upside down. We tend to imagine a computer as coldly competent — brilliant at the technical, hopeless at the tender. The study found almost the reverse.

On the general skills of therapy — expressing empathy, validating a person’s feelings, building a sense of working together — the chatbot did not merely hold its own. It outperformed the human practitioners. But on the specific, technical craft of CBT — pinpointing the key belief beneath a worry, guiding someone towards their own insight, tailoring the approach to the particular person in the chair — it faltered. The machine was warm but not always wise: a generous listener that struggled to do the precise, individual work that lifts a conversation into genuine treatment.

The overall numbers tell a measured tale. A score of 40 out of a possible 66 is the accepted threshold for adequate clinical competence. The chatbot reached it in just 30 of the 65 sessions — better than a coin toss, but strikingly inconsistent from one conversation to the next. Against the full pool of human therapists, who averaged 40.3 points, the chatbot’s adjusted score of 38.1 sat a little lower. Yet when the researchers compared it only against the six most rigorously conducted human studies, the gap vanished: the difference was no longer statistically meaningful.

It was that inconsistency, more than any average, that gave the researchers pause. “We were surprised by how much variation the LLM-chatbot showed in its skillfulness across CBT sessions,” Herbener told the science outlet PsyPost. The same system could be quietly excellent in one conversation and fall short in the next — and we do not yet fully understand why.

An honest word about the limits

It would be easy to read a headline here and conclude that the robot therapist has arrived. It has not, and the researchers are refreshingly candid about why. This was a single session with young adults in mild distress — a world away from the complexity of someone in real crisis, where the demands on a therapist’s judgement rise sharply. One half-hour cannot capture the slow, patient architecture of a full course of treatment: the homework reviewed, the trust built week by week, the relationship that so often does the real healing.

There are subtler cautions too. The rating scale was designed for video or audio, where a rater can hear a catch in the voice or read a face; applied to plain text transcripts, some of that human nuance may simply be invisible. The raters also knew they were judging a machine rather than a person, which can quietly colour a score. And large language models carry a peculiar habit the researchers name sycophancy — a tendency to over-agree, to soothe rather than to challenge. Yet gently challenging an unhelpful belief is often exactly what helps a person change. A therapist who only ever nods is not, in the end, doing therapy. As Herbener puts it, good observable skill “is not the same as clinical effectiveness.”

So what does this mean for you?

Most of us are not commissioning a clinical trial; we are simply deciding whether to open an app when the mind feels loud. The research does not forbid that — but it does hand us a wiser way to go about it. It helps to think of an AI chatbot not as a therapist, but as a well-read, endlessly patient companion: wonderful for a first step, no substitute for the real thing. Here are a few practical ways to lean on these tools kindly, and safely:

  • Treat it as a bridge, not a destination. A chatbot can help you name a feeling or steady a racing three-in-the-morning mind, but it is a waiting-room companion, not the appointment itself.
  • Keep a human in the loop. If distress is deepening, take what you have written to your GP, to NHS Talking Therapies (which, in England, you can refer yourself to without a doctor’s note), or to a person you trust. A machine should widen your circle of support, never replace it.
  • Notice the flattery. If the chatbot only ever agrees with you, gently push back on its behalf. Ask it, “What might I be getting wrong here?” — and weigh the answer with your own judgement.
  • Guard your privacy. Share as you would with a stranger who keeps notes, and think twice before pouring in details you would not want stored somewhere.
  • Know where the edge is. For thoughts of self-harm or genuine crisis, step away from the app and reach a person. In the UK, you can call Samaritans free on 116 123, at any hour of any day.

The kindness in the machine — and the wisdom in us

Perhaps the most human finding in all of this is the one hiding in plain sight. The machine was at its best when it was warm, and at its weakest when it needed to truly understand one particular person. That should not frighten us; it should hearten us. It tells us that the hardest, most valuable thing in care is not information, which computers have in abundance, but attunement — the slow art of meeting someone exactly where they are. That remains, for now, gloriously, stubbornly human.

So if a chatbot has helped you put words to a hard night, be glad of it. Then take the next step it cannot take for you. Reach out to a real person today — a friend, a GP, a helpline — and let the warm voice of the machine be the thing that gave you the courage to go and find a warmer one still.


Source: Herbener, A. B., Zachariae, R., Klincewicz, M., Thøgersen, M. B., Hermann, M. R., & Damholdt, M. F. “Exploring the therapeutic competencies of large language models: Observational study and comparison with meta-analytical estimates for human therapists,” Computers in Human Behavior: Artificial Humans (2026). Reported by Eric W. Dolan for PsyPost, 24 August 2026.

This piece discusses mental health, including a brief mention of self-harm and crisis. If you are struggling yourself, you are not alone and support is available — in the UK you can call Samaritans free on 116 123 at any time.


Discover more from Mental understanding

Subscribe to get the latest posts sent to your email.

Leave a comment