A human peer and an AI friend can occupy the same screen. Both can reply, ask a question and keep a conversation going. In a two-week randomised trial with first-year university students, however, only the human relationship changed the broader loneliness score.
Ruo-Ning Li and colleagues assigned students to text a randomly paired peer, message a supportive chatbot called Sam, or write a one-sentence daily journal. The students who spoke to another person reported less loneliness at the end. Those assigned to Sam did not differ from the journal group.
This is not treatment advice. I am reading one study carefully and asking what it does, and does not, show.
The finding matters because Sam was not a generic bot left to improvise a pleasant manner. It had memory and a prompt deliberately built around active listening, empathy, understanding and validation. Yet the outcome also carries a useful complication: the AI reduced negative mood even though it did not reduce loneliness.
Comfort and connection were not the same result.
How the trial worked, and which sample each number describes
The peer-reviewed paper in the Journal of Experimental Social Psychology reports that 306 students completed the initial survey and were assigned to a condition. Three withdrew and six were excluded because they completed fewer than three daily surveys.
That left 296 people in the daily dataset, contributing 3,703 surveys. The primary before-and-after loneliness analysis was narrower: 276 completed the relevant measures and one lacked a baseline score, leaving 275. The average participant was 18, 72 per cent were female, and recruitment took place during the first semester at a Canadian university.
The headline’s 296 refers to the final daily-survey sample; the central pre-post comparison rests on 275.
The hypotheses, methods and analysis plan were preregistered. The authors also posted the data, materials and analysis code on OSF.
For 14 days, participants used private rooms on Discord. Those in the human and AI groups were asked to send at least one meaningful message a day, defined as more than a greeting. The control group wrote a one-sentence summary of the day. At 9pm, anyone who had not yet participated received a text reminder.
This was not conversation versus silence. Every group had a daily ritual, a private place to put experience into words, repeated surveys and a prompt to return.
Sam ran on ChatGPT-4o mini and had memory for continuing, personalised conversations. The researchers told it to act as a friendly, positive and supportive AI friend. Its instructions drew on relationship science, including active listening, understanding and validation.
The human participants were randomly paired with other first-year students. They met briefly in person during the first laboratory visit before continuing by text, a detail that becomes important when interpreting the comparison.
Only the human group moved the loneliness measure
The researchers used the 20-item UCLA Loneliness Scale. Participants rated statements such as lacking companionship from one to four, with reference to the previous two weeks. Baseline loneliness did not differ significantly across groups.
After the model controlled for those baseline scores, adjusted post-study loneliness averaged 1.85 in the human condition, 1.98 in the AI condition and 2.00 in the journal condition. The human group was significantly lower than both alternatives. The AI and control groups were statistically indistinguishable.
The raw within-group pattern tells the same story. Human participants fell from an average of 2.04 to 1.86, a change of 0.18 points. The AI and journal groups each declined by about 0.04 points, neither a significant shift.
This is one study, not settled consensus. It does not show that Sam harmed students or increased loneliness. It shows that, in this design and over this fortnight, the bot did no better on the main outcome than writing one sentence about the day.
The bot was used, and it did improve one kind of feeling
The message logs rule out a simple explanation that students barely engaged with the machine. They sent an average of 8.95 messages a day to the AI and 10.23 to a human, a difference that was not significant. They also wrote more words to Sam, averaging 81.61 a day versus 65.45 to a peer.
The exploratory mood measures complicate any verdict that the AI was useless. Both conversation groups reported less negative mood than the journal group, in daily ratings and at the end. Only the human group showed higher positive mood. A daily question about social connection during the Discord exchange found no significant group differences.
Silicon Canals recently examined why feeling heard can make AI companionship subjectively meaningful. That article drew on experiments in which AI companions produced short-term reductions in loneliness, partly through perceived listening. The underlying Journal of Consumer Research programme measured immediate or short-horizon effects, while Li’s trial asked whether two weeks of repeated use altered a broader assessment of loneliness.
The results can coexist. A conversation may take the edge off a bad evening without changing someone’s sense of whether their social needs are met.
Mutuality is a plausible mechanism, not a demonstrated one
The most revealing evidence came from exploratory analysis of the conversations. A language model rated 2,045 daily exchanges for empathy and engagement, with reliability checked against a smaller human-coded subset. Sam received the highest empathy and engagement ratings overall. Still, the students themselves expressed more empathy and put more effort into the exchange when the partner was human.
The researchers suggest that receiving care may be only half of what helps a relationship matter. A person can also be needed, choose to respond and offer attention. A message from another student during midterms may carry evidence that the person chose to make time.
Behaviour after the study is consistent with that reading. During an optional extra week, 33 per cent of the human group continued chatting, compared with 14 per cent of the AI group and 3 per cent of the journal group. Among human pairs, 37 per cent exchanged contact details.
That is continued engagement, not proof of the mechanism. The trial was not powered to test mediation, and closeness ratings did not differ between human and AI partners. The brief face-to-face meeting also gave the human pairs an advantage.
The result does sit naturally beside earlier Silicon Canals reporting on how strangers underestimate the interest others will show in deeper conversation. The shared point is modest: an unfamiliar person is not an empty condition. There is another set of needs, choices and future possibilities on the other side.
What a two-week student trial cannot decide
Most participants were not highly lonely at the beginning. The modal student said they rarely experienced loneliness. They also shared a campus transition and social opportunities. The result may not generalise to people whose available human contact is scarce.
Fourteen days is short. The study tested one model, one prompt and a modest level of use. Participants were paid $20 or received course credit, and reminders created consistency that an ordinary product cannot assume. The study had 80 per cent power to detect a moderate between-group effect, equivalent to a standardised difference of 0.40, so a smaller benefit could have gone undetected.
The main outcome was self-reported. The established scale does not show what happened to each participant’s offline relationships, and groups did not differ significantly in reported new friends.
There is no general contest here between all humans and all chatbots. Another system, timeframe or population could produce a different result. The authors propose testing chatbots that help people improve contact with others rather than serve as surrogate companions.
That idea connects with my earlier look at how close friendship accumulates through repeated time together. A product might make those hours easier to arrange. It cannot assume that simulating the finished relationship produces the same outcome.
The claim a product can honestly make is narrower
An AI companion does not need feelings of its own to help someone feel better during an exchange. Sam appears to have done that. The trial nevertheless gives developers, universities and users a reason to separate immediate emotional relief from a durable reduction in loneliness.
Peer-matching is not effortless or risk-free. It requires consent, moderation, privacy safeguards and a plan for bad pairings. This study does not provide a ready-made campus programme. It does show that a low-tech human intervention outperformed a deliberately supportive chatbot on the outcome the trial set out to test.
For these students, the difference was not how much the partner typed or how empathic the replies appeared. It was that only one partner could choose to care, receive care in return and still be there after the research room closed.