A group of four people is asked to solve a mystery together. Three of them already know each other. The fourth is an outsider, someone who does not share the group’s usual social ties. Intuition suggests the group with the outsider should feel more awkward, less cohesive, and probably perform worse. In a controlled experiment, it performed better, and felt worse doing it.
What the actual study measured
The experiment comes from a 2009 study in Personality and Social Psychology Bulletin, by Katherine Phillips, Katie Liljenquist, and Margaret Neale, then at Northwestern and Stanford. Four-person groups worked through a murder-mystery style decision task with a correct solution, and the researchers varied whether the group’s newest member belonged to the same social category as the rest, or was socially distinct from them. Groups that included a socially distinct newcomer reached the correct answer more often than groups where everyone belonged to the same in-group.
The surprising part was not the performance gain. It was what came with it. According to a summary of the research published by Kellogg School of Management, where Phillips held a faculty position, members of the more accurate groups reported feeling less confident in their decision and rated their own group process as less effective, even as they objectively performed better. The people who did the better work did not feel like they had done better work.
Where the benefit actually came from
This is one study, not a settled account of exactly how cognitive diversity operates in every setting, and it examined social distinctiveness specifically, rather than every dimension groups might vary on. But the mechanism it points to is worth sitting with. The performance gain in Phillips, Liljenquist, and Neale’s data did not come from the newcomer supplying some special insight or piece of information the others lacked. The groups performed better even when the newcomer’s actual opinion happened to be wrong.
What changed was how the existing members behaved once someone socially distinct was in the room. According to the study, oldtimers anticipated more disagreement from a newcomer who did not share their background, and that anticipation made them work harder: rehearsing their reasoning more carefully, considering alternatives they might otherwise have skipped, and treating the discussion less like a formality among people who already agreed. The social friction of not being sure whether a new person shared your assumptions appears to have done the cognitive work, not any particular idea the newcomer brought with them.
A separate, more theoretical version of the same claim
A related and much more widely cited idea comes from a different kind of research entirely. In a 2004 paper in the Proceedings of the National Academy of Sciences, the researchers Lu Hong and Scott Page built a mathematical model of problem-solving agents and found that, under the conditions their model specified, a randomly selected group of diverse problem solvers could outperform a group made up of the individually highest-scoring solvers. This became known as the diversity-beats-ability result, and it is frequently cited well beyond the narrow computational context it came from.
The result is a model, not a study of actual human groups, and later researchers have tested how far it generalizes. A 2022 replication published in ReScience C, by the researcher Lukas Wallrich, reran Hong and Page’s original model and confirmed the core finding held under the conditions originally tested, while other researchers examining more varied problem landscapes have found the effect is narrower and more condition-dependent than the original framing suggested. The honest version of the claim is that diversity can beat ability inside a specific mathematical setup, not that it reliably does so in every group, every task, and every landscape of possible solutions.
Why the discomfort might be the point, not the cost
Put next to each other, these two lines of research suggest something different from the usual pitch for inclusive thinking, which tends to promise that a wider range of perspectives feels enriching and produces better ideas as a direct result. The Phillips, Liljenquist, and Neale data suggests the mechanism can run the opposite way: the benefit shows up partly because working with someone whose assumptions you cannot take for granted is mildly uncomfortable, and that discomfort pushes a group to actually do the reasoning it might otherwise skip when everyone already agrees.
None of this suggests that groups should manufacture social distance for its own sake, or that comfort and good decisions are opposed in some general way. It suggests something narrower: that the felt experience of a group discussion, whether it seems smooth and confident or effortful and uncertain, is not a reliable signal of how well the group is actually reasoning through a problem. A group that feels settled and in agreement may simply be skipping steps that a slightly less comfortable group is forced to take.
What this does not settle
This is a small number of studies, using specific tasks and specific definitions of social distinctiveness or diversity, not a general law covering every team, every industry, and every kind of difference a group might contain. The research does not say discomfort itself causes better decisions, only that in these particular experiments, the presence of someone whose views could not be assumed in advance changed how carefully a group worked. Whether that same mechanism operates the same way for other kinds of difference, over longer stretches of time than a single lab task, remains a more open question than either the original theorem or the newer studies fully answer.
What this suggests for how a team reads its own comfort
For a team evaluating its own decision-making process, the practical implication of the Phillips, Liljenquist, and Neale finding is an awkward one: a meeting that feels efficient, in agreement, and pleasant is not obviously evidence that the group reasoned well, and a meeting that feels effortful or slightly uncertain is not obviously evidence that it reasoned badly. A hiring or team-composition decision built purely on who will make a group feel most comfortable together risks selecting for exactly the dynamic this research associates with weaker collective reasoning, even though comfort is what most people would name if asked what a good team meeting feels like.
None of this is a case for treating friction as a goal in itself, or for assuming that any group containing disagreement is therefore reasoning well. The research described here measured accuracy on a specific task with a correct answer, not general team morale or long-term retention, both of which matter for reasons this body of research was never designed to speak to.