University students who used an AI chatbot for 12 weeks ended with anxiety scores about 2 points lower than students in weekly group therapy, in a randomized trial of 977 students in Israel.1 The gap was still there 3 months later. But it came less from the chatbot users improving a lot than from anxiety creeping up in everyone else, including the students in group therapy.
Research Highlights
- Anxiety favored the chatbot: After 12 weeks, chatbot users scored 2.1 points lower than group-therapy members on a 21-point anxiety scale, and 1.6 points lower at 3 months.1
- Much of the gap came from rising anxiety in the comparison arms: Average anxiety fell 0.9 points with the chatbot but rose about 1.2 points with both group therapy and no treatment.1
- Depression and PTSD did not follow suit: The chatbot beat no treatment on depression but not group therapy, and no arm separated on PTSD symptoms.1
- Lonely students used it twice as much: Highly lonely users spent about 124 minutes a week with the chatbot vs. 61 for the least lonely, and their anxiety dropped more.1
- A company-linked, unblinded trial: The lead author consults for the chatbot’s maker and holds stock options, and every outcome was self-reported.1
Most chatbot trials compare an app against nothing, a waitlist, or a reading assignment. That shows whether the chatbot beats doing nothing, not whether it can stand in for help from a human.
Shoshani et al. added a human-led arm: 12 weeks of group therapy run by licensed clinical psychologists. That turns the trial into a large, direct test of a therapy chatbot against active, human-delivered treatment.1
What the AI Chatbot and the Group Therapy Actually Were
The AI arm used Kai, a commercial emotional-support system from Kai.AI that runs inside WhatsApp or iMessage. It is not a general-purpose chatbot like ChatGPT. The paper describes separate components for detecting emotion, classifying what the user wants, running structured exercises, and generating open conversation, built with fine-tuning, controlled prompts, and a curated library of therapy exercises.1
In practice, students got:1
- A personalized plan: an intake questionnaire set a mix of exercises drawn from CBT, acceptance and commitment therapy (ACT), dialectical behavior therapy (DBT), mindfulness, and positive psychology.
- Daily routines: morning prompts, evening reflection check-ins, reminders, weekly progress summaries, and an auto-generated journal.
- Open chat: free-form conversation at any hour, with a target of at least 3 sessions a week.
The group therapy arm met in person for 12 weekly 90-minute sessions, about 20 students per group, each led by a licensed clinical psychologist. The program was “integrative,” meaning it combined several approaches: CBT-style thought challenging, ACT and DBT skills, relaxation and breathing, gratitude exercises, and open sharing among group members.1
A third group was a waitlist: no intervention for the trial period, then access to either program afterward.
Both programs drew on similar therapy toolkits. The biggest differences were in delivery: a phone available around the clock and tailored to one person vs. a fixed weekly meeting with a large group and a commute.
Who was in the trial
- Participants: 977 university students, ages 18 to 32, average 23, about half women.1
- Symptoms: mild on average. Starting scores were 7.9 on the PHQ-9 depression scale and 7.4 on the GAD-7 anxiety scale, both in the “mild” band, though about 1 standard deviation above typical young adults.
- Excluded: anyone with psychosis, acute suicidality, severe psychiatric illness, or current psychotherapy or psychiatric medication.
- Randomization: 329 to the chatbot, 324 to group therapy, and 324 to the waitlist. Students had to be willing to do either program and received course credit.
Anxiety Fell With the Chatbot and Rose in Both Other Arms
Anxiety was measured with the GAD-7, a 7-question survey scored 0 to 21, where 5 to 9 counts as mild and 10 or more usually prompts a closer clinical look.1
After 12 weeks, chatbot users scored 2.10 points lower than the group-therapy arm (95% CI 1.60 to 2.59 points) and 2.08 points lower than the waitlist. Group therapy and the waitlist were indistinguishable (p = 1.00). At the 3-month follow-up, the chatbot was still ahead of group therapy by 1.58 points and of the waitlist by 2.01.1

The chart shows where the 2-point gap came from. Chatbot users improved by 0.9 points, a small change (within-group d = 0.20). Group-therapy members and waitlisted students both got about 1.2 points worse.1
The paper does not explain why anxiety rose in both comparison arms. The authors do note that people in Israel live with chronic exposure to sociopolitical conflict, which may shape distress levels; data collection began in April 2025. Whatever the cause, the chatbot’s advantage was mostly that its users did not drift upward.
The clinical-threshold results look more dramatic. Among students who started with an anxiety score of 10 or higher, the share who dropped below 10 by week 12 was:1
- Chatbot: 52 of 92 (56.5%)
- Group therapy: 10 of 96 (10.4%)
- Waitlist: 8 of 110 (7.3%)
At follow-up the chatbot still led, 50.9% vs. 20.6% and 14.5%, though by then about a third of each arm had stopped answering surveys. Crossing a cutoff can also mean moving from 10 to 9, so those percentages overstate how many students felt a large change.
Depression, PTSD, and Well-Being Were Not Uniform Wins
The pattern shifted outcome by outcome:1
- Depression (PHQ-9, 0–27): the chatbot beat the waitlist by 1.91 points and group therapy beat the waitlist by 1.25. The chatbot’s 0.65-point edge over group therapy (p = .046) did not clear the paper’s own stricter significance threshold of p < .01, and the 0.86-point gap at follow-up was not significant either (p = .14).
- PTSD symptoms: no arm differed (p = .686). Scores dipped slightly in all 3 arms, including the waitlist.
- Well-being (WHO-5): the chatbot beat group therapy by 5.6 points after treatment and 6.3 at follow-up. Group therapy also edged out the waitlist (p = .031 after treatment).
- Life satisfaction: the chatbot’s lead over group therapy was 1.0 points after treatment (p = .036, again above the paper’s threshold) and 2.1 points at follow-up (p = .001).
For scale, a drop of 5 or more PHQ-9 points is a commonly used benchmark for a clinically meaningful change.2 The chatbot arm’s average depression score fell 0.9 points. These were mildly distressed students, so there was little room for large drops, but the averages describe modest shifts, not recoveries.
Why “Beat Group Therapy” Needs Context
The trial does show that this chatbot outperformed this group program on anxiety and well-being. Several design features limit how far that result travels:
- Not a test of therapy in general. The comparator was a brief university program with about 20 people per group. The trial says nothing about one-on-one therapy, or about longer or more intensive group treatment for people with clinical diagnoses.
- Delivery was not matched. The chatbot was in students’ pockets every day and tailored to each person. Group therapy was one scheduled meeting a week, and some members described commuting and time as barriers. The authors acknowledge that the gap may reflect access and personalization rather than the therapy content itself.1
- Total time was similar. Chatbot users averaged 82 minutes a week; group members attended 10.3 of 12 sessions. By simple multiplication, that is roughly 16 hours vs. 15.5 hours over 12 weeks, so the chatbot did not win simply by getting more contact time.
- Everyone knew their arm. Students could not be blinded, and every outcome was a self-report questionnaire. Expectations about a new AI tool can move those scores.
- Company ties. The first author is a paid consultant for Kai AI with stock options, a second author is a Kai employee, and there was no outside funding.1
- Registration timing. The paper calls the trial pre-registered, but its own dates put registration on May 12, 2025, about a month after data collection began on April 10.
- Early-access version. npj Digital Medicine posted the paper on July 7, 2026, as an unedited accepted manuscript, so details could still change.
No equivalence or noninferiority test was run, which is the kind of design needed to claim that 2 treatments work about equally well. So the most accurate summary is narrow: for mildly distressed students who were not in treatment, this chatbot reduced anxiety relative to both a large, brief group program and no treatment.
Lonely and Insecurely Attached Students Benefited More
The paper’s title question was who benefits. Researchers measured 3 relational traits at baseline:
- Loneliness: feeling short of meaningful connection.
- Perceived social support: believing there are people you can rely on.
- Attachment style: anxious attachment means worrying about rejection and needing reassurance; avoidant attachment means keeping emotional distance and prizing self-reliance.
Loneliness mattered most. The chatbot’s estimated anxiety benefit over the comparison arms was 2.92 GAD-7 points at high loneliness vs. 1.14 at low loneliness. Loneliness also strengthened its effects on depression (2.11 vs. 1.23 points) and well-being.1
For anxiety only, the chatbot also worked better for people with high anxious attachment (2.60 vs. 1.55 points), high avoidant attachment (2.41 vs. 1.99), and low social support (2.65 vs. 2.11). The avoidant result ran against expectations that avoidant people would distrust AI. The authors suggest that the anonymity of a chatbot may lower the bar for people who keep others at arm’s length.1
The link runs through how much people used it
The lonelier students used Kai far more. The 51 most lonely wrote about 2,490 words a week and spent 124 minutes with it, vs. 1,217 words and 61 minutes for the 71 least lonely.1
Within the chatbot arm, the number of messages sent predicted anxiety reduction. Once message count was accounted for, loneliness and attachment no longer predicted outcomes on their own. The researchers read this as engagement carrying the effect.1
That is a correlation, not proof that more chatting causes more relief. Students who were already starting to feel better may have engaged more, or the heavy users may have differed in other ways. The mediation model was added after the fact, was not preregistered, and had only borderline statistical fit.
Two other findings complicate the “AI for the lonely” story:
- Group therapy helped lonely students too. In an exploratory analysis, lonelier students in group therapy also had larger anxiety reductions. Their well-being gains were smaller, though.1
- Heavy use has looked worse elsewhere. In a 4-week MIT Media Lab and OpenAI randomized study of 981 adults using a general-purpose chatbot, people who voluntarily chatted more had consistently worse outcomes on loneliness, real-world socializing, emotional dependence, and problematic use. That study has been posted as a preprint and not yet peer reviewed.3
The contrast may reflect the tools. Kai was built around therapy exercises and included prompts to notice real-world relationships, while the MIT/OpenAI study tested a general-purpose chatbot under several text, voice, and conversation-topic settings. Neither study can say whether heavy use by lonely people builds skills or replaces human contact over months or years. Kai’s trial ended 3 months after treatment.
How the Result Compares With Other Chatbot Trials
Typical effects are small. A 2026 meta-analysis of 48 randomized trials and 28,071 participants found that conversational agents reduced depression by a standardized mean difference of 0.27 and anxiety by 0.20.4
A standardized mean difference expresses the gap between groups in standard deviations, so 0.2 is small. Kai’s 2.1-point anxiety gap works out to roughly 0.4 standard deviations, on the larger side of that literature, though much of it came from comparison arms worsening.
Earlier student data pointed the other way on anxiety. In a 2017 trial of Woebot, a scripted CBT chatbot, 70 students were randomized to use it for 2 weeks or to read a depression ebook from the National Institute of Mental Health.
Woebot reduced depression more than the ebook (p = .01), but anxiety fell in both groups without a clear difference.5 One of its authors worked for Woebot’s developer, a conflict of interest similar to the one in the Kai trial.
Sicker samples showed bigger changes. In the 2025 Therabot trial, 210 adults with clinical-level depression, anxiety, or eating-disorder risk used a generative AI chatbot for 4 weeks or waited.
PHQ-9 depression scores fell 6.1 points with Therabot vs. 2.6 on the waitlist, with large effect sizes (d = 0.85 to 0.90 for depression, 0.79 to 0.84 for anxiety).2 Therabot’s comparator was a waitlist, and its authors called for trials against active treatments like the one Kai faced.
Read together, chatbots produce small effects against weak comparators in mild samples, larger effects in more symptomatic people, and, in the Kai trial, an anxiety advantage over a human-led group. There is still no large trial against individual therapy.
How Kai Handled Crisis Risk
Safety is the reason many clinicians hesitate over therapy chatbots. The trial excluded anyone acutely suicidal, and Kai layered several protections:1
- An algorithm screened messages in real time for signs of acute distress, such as suicidal thoughts or severe hopelessness, and automatically sent local crisis hotline and referral information.
- Content filters and refusal rules blocked requests for diagnoses or clinical direction.
- Licensed clinicians reviewed flagged high-risk conversations in real time and could reach out directly.
The paper reports no numbers for any of this: not how many conversations were flagged, how often clinicians intervened, or whether any adverse events occurred.
The Therabot trial offers a benchmark. Clinicians reviewed every reply after it was sent and stepped in 15 times over safety concerns such as suicidal ideation and 13 times to correct inappropriate responses, among 106 users in 4 weeks.2
With 329 users over 12 weeks, the Kai arm likely produced at least some flagged conversations, and leaving them out of the report is a real gap.
What Students Can Take From the Trial
- For mild, everyday anxiety: a well-designed, therapy-focused chatbot with human safety monitoring is a reasonable thing to try, especially when counseling has a waitlist.
- For depression, trauma symptoms, or anything severe: this trial offers little. The chatbot did not beat group therapy on depression, did nothing specific for PTSD symptoms, and excluded anyone at acute risk.
- If loneliness is part of the problem: the lonely students used Kai heavily and did well on anxiety. It is still worth watching whether the chatbot is becoming a replacement for people rather than a bridge back to them.
- Not all chatbots are equal: Kai had structured exercises and clinician review of flagged risk. A general-purpose assistant has neither by default.
Questions About AI Chatbots vs. Group Therapy
Did the AI chatbot work better than therapy?
The chatbot outperformed the brief group program tested, which had about 20 students per group, on anxiety and well-being. Life satisfaction separated at the 3-month follow-up, but not at 12 weeks under the trial’s corrected significance cutoff. It did not clearly beat the group program on depression. The trial did not test individual therapy.1
Why did group therapy do no better than the waitlist on anxiety?
The paper does not say. Anxiety rose by about the same amount in both arms. Possible factors include the large group size, a once-weekly schedule during a stressful period, and questionnaire-based outcomes, but none of these was tested. Group therapy did beat the waitlist on depression right after treatment and on well-being.1
Does using a chatbot more make it work better?
Within the Kai arm, heavier users improved more on anxiety, but that link is correlational. In the MIT/OpenAI study of a general-purpose chatbot, heavier voluntary use went with worse loneliness and dependence outcomes.1, 3
How big were the improvements for chatbot users?
Small in absolute terms. Average anxiety fell from 7.2 to 6.3 on a 21-point scale, and depression from 7.6 to 6.7 on a 27-point scale. Roughly half of those who started above the clinical cutoff of 10 ended below it.1
References
- Attachment, loneliness, and social support as moderators of conversational AI–based mental health outcomes. Shoshani A, et al. npj Digit Med. 2026 (unedited accepted manuscript, published online July 7, 2026). doi:10.1038/s41746-026-02974-y
- Randomized Trial of a Generative AI Chatbot for Mental Health Treatment. Heinz MV, et al. NEJM AI. 2025;2(4). doi:10.1056/AIoa2400802
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study. Fang CM, et al. arXiv preprint. 2025 (revised October 2, 2025). doi:10.48550/arXiv.2503.17473
- Effectiveness of AI and rule-based conversational agents for depression, anxiety and stress: A meta-analysis. Mokhtari Masoumi Alamdarloo S, et al. npj Digit Med. 2026;9:650. doi:10.1038/s41746-026-02820-1
- Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial. Fitzpatrick KK, et al. JMIR Ment Health. 2017;4(2):e19. doi:10.2196/mental.7785