You have probably used the polite version. Someone disagrees with you about something that matters, and before you explain why they are wrong, you tell them you are sure they want what is best. It costs nothing, and it lowers the temperature.

The plainer alternative is to skip the compliment: name the reasons you think they hold, and call them valid. A preregistered experiment published on 12 August in Communications Psychology ran the two against each other. Across four measures of how the encounter felt, the plainer version produced effects roughly twice the size. Neither framing shifted opinion any further than the bare counter-argument did on its own.

The two sentences the experiment turns on

Li Li Hwangpo, Lindi R. Shepard, Lisa Nehring, Nan Mu, Katherine J. Cornwall and Hunter Gehlbach — four at the Johns Hopkins School of Education, one at Indiana University Bloomington, one at the University of Wisconsin–Madison School of Medicine and Public Health — recruited 579 US adults on Prolific on 17 May 2024, and registered the hypotheses on the same date, which is all the paper establishes about the order of the two. Twenty-one were dropped for answering twelve consecutive questions identically, leaving 558. The survey took about ten minutes and paid $2.41.

Participants read a short brief on climate change education policy, including a line telling them New Jersey was the first state to require it across grades and content areas, then said whether they supported a national mandate for it in K–12 public schools and how strongly. Only after that were they randomized, within their own opinion group, into one of three conditions.

Everyone then read a mock Facebook post from a fictional public-school teacher called Taylor Harris — gender-neutral name, gender-neutral pronouns, no other biography, so there was, in the paper’s phrase, “limited chance for participants to identify preexisting commonalities with the teacher.” The post argued against whatever the participant had just said. In the control condition, that was all it did.

The teacher never asked anybody anything. That is the design’s whole point: the arguments being affirmed are guesses. In one treatment condition, the post set out the arguments a person on the participant’s side might hold, called them sound, and moved on with the phrase “despite these very valid reasons…” or “despite these very real concerns…” depending on which side the reader was on. In the other, the teacher named the participant’s likely motives and endorsed them: “while I appreciate these intentions and our shared motivations…”. Everything after that was counter-argument.

Where the two versions separated

Both treatments worked on the relational measures, and one worked harder. On whether the teacher seemed to have taken the reader’s perspective, the argument version produced a covariate-adjusted Cohen’s d of 1.12 against control (95% CI 0.88 to 1.34); the intention version, 0.45 (0.26 to 0.67). On perceived similarity: 0.66 against 0.33. Expected relationship with the teacher came in at 0.69 against 0.36. On whether the teacher and the information seemed fair: 1.37 against 0.56, the largest effect anywhere in the study.

In the results the authors write that “across these four outcomes, the argument-affirming SPT treatment produced effects that were approximately twice those of the intention-affirming SPT treatment,” a claim their conclusion narrows to “for most outcomes” and the abstract hedges again as what “estimated effect sizes further suggested.” They also say the direction surprised them: the team “had tentatively expected the opposite,” on the theory that validating someone’s deeper intentions would reach further into who they are.

Two results did not hold up. A fifth relational outcome, whether readers reciprocated by working harder to understand the teacher, was significant for the argument condition at p = .017, but the paper reports that its confidence interval included zero once a Holm-Bonferroni correction for twelve tests was applied; the intention condition was not significant to begin with (p = .220). The authors record three deviations from their preregistered analysis, all of them additions rather than substitutions: the Holm-Bonferroni correction was applied as what they call an additional robustness check, and an opinion-change model without the initial-opinion covariate as an exploratory sensitivity check. Under the preregistered analysis on its own, the argument arm’s interval on that fifth outcome excludes zero.

And on the outcome most people would care about, the treatments added nothing detectable. Everyone moved a little. Across the whole sample, control group included, opinions shifted 0.80 points toward the opposing side on the ten-point scale. But the gap between treatment and control was small and non-significant: p = .195 for the argument condition, p = .602 for intentions, and similar in a sensitivity model that dropped the initial-opinion covariate (p = .192 and p = .564). The effect sizes are d = 0.13, 95% CI -0.05 to 0.33 for the argument condition and d = 0.05, -0.15 to 0.25 for intentions, and the paper reports no equivalence test, so this is an undetected difference rather than a demonstrated absence. A counter-argument on its own already did whatever moving there was to do.

The fairness effect is the largest number in the study, and part of it is built into the stimulus. Argument-condition readers who started out supporting the mandate estimated that 26.99% of the post’s information favored their side, against 8.40% among control-group supporters. What they were estimating is their own perception, and by design the treatment posts did carry more of their position, because the affirmation was bolted to the front of them. Calling a two-sided message fairer than a one-sided one describes what the reader was handed. Nothing in the paper’s limitations section raises the possibility. That reading is ours, not the authors’.

The four relational outcomes are also less independent than four numbers suggest. The authors report intercorrelations between them of r = .71 to .85 and raise it themselves, though they defend the measures first: “Although these constructs are theoretically distinct and have reasonable internal consistency (Cronbach’s ɑ = .86 to .91), a broader range of distinct outcomes might benefit future studies.” Nor is the two-to-one gap formally tested anywhere. It is a comparison of point estimates between two treatment arms, and the individual ratios run from about 1.9 to about 2.5 rather than landing on two. On two of the four outcomes, perceived similarity and expected relationship, the two arms’ intervals against control overlap.

One more thing the measure cannot see. The perceived-perspective-taking scale averages four items — effort, motivation, clarity and accuracy — into a single score, so the first and the last cannot be pulled apart. The authors say so directly: the scale “did not distinguish perceived SPT accuracy from perceived SPT effort,” and individuals “might feel that a perceiver is trying hard to take their perspective but doing so inaccurately.” The paper’s founding anecdote concedes the same gap. Environmental campaigners opened town halls in coal country by praising coal, efforts that the paper says “appeared to have been appreciated,” and in the same breath it notes that “merely trying to take the coal miners’ perspectives did not ensure accurate SPT.”

A ten-minute argument with a stranger who does not exist

This is a single experiment on a single issue with a US sample, and its authors say so. The encounter, they write, was “text-based, impersonal, and unidirectional, potentially making the stakes for taking the perspective of the teacher fairly low,” with none of the facial expression, gesture or tone a real argument carries; the “U.S. political context is unique, potentially affecting the international and cultural generalizability of our findings.” The paper also arrived with the publisher’s early-access banner attached, warning that this is an unedited manuscript in which “there may be errors present which affect the content.”

The sample leaned one way, too. On a ten-point scale from strong opposition to strong support for the mandate, the average starting position was 7.54, and participants rated themselves slightly liberal on average. Nothing here was measured on people whose disagreement was with someone they have to see at Thanksgiving.

That distinction matters more than the effect sizes do. Nobody in this experiment was arguing with someone they love. This piece is a reading of one ten-minute study, written by people with no clinical training, and it is not advice about anyone’s relationship. A disagreement that is corroding a marriage or a family is work for a couples or family therapist, and no choice of opening sentence substitutes for one.

The advice the paper gives writers

The paper’s practical-implications section addresses advocacy groups, teachers and, specifically, “bloggers, op-ed writers, journalists, and others addressing broad, ideologically diverse audiences,” who “might begin by acknowledging the arguments or intentions of those who disagree.” It is the least guarded prose in the article — the gesture “can produce a powerful bridge,” such efforts “are likely to enhance audience perceptions of the communicator” — and it runs ahead of a design in which nobody spoke to anybody, nothing was measured after the ten minutes were up, and the bridge in question is a self-reported score about a person who does not exist.

The narrower claim survives, and it is worth having. If you want the person you are about to contradict to think better of you afterward, the evidence here favors making a visible attempt to state their case before you take it apart. Bringing them round any faster is not on offer: the paper reports “small, non-significant differences” between treatment and control on opinion change. The polite version you reach for, the assurance that you know they mean well, came second in the one setting this experiment tested it.