One of the cases the paper opens with is hormone therapy. Analyses from the Nurses’ Health Study linked postmenopausal estrogen therapy to reduced cardiovascular risk; randomized trials later showed the opposite, with benefits limited to younger women. Nobody had faked a number. The data had been read as an answer to a question it was never built to answer.
A team at the University of Pennsylvania has now put a size on how often that reading gets invited. In a paper published in Nature Human Behaviour on 24 August 2026, Calvin Isch, Timothy Dörr, Neil Fasching, Grace Jennings and Duncan J. Watts classified 194,631 social science articles that relied purely on cross-sectional data — a single snapshot, no time ordering, no randomization — and coded 46.3 percent of them as using causal language about their own results in the title or abstract. The rate held near 20 percent from 1980 to 2000, then climbed to over 60 percent by 2024.
Nearly half, in a very particular pile of papers
The scope matters more than the headline number. The corpus was assembled by searching five ProQuest databases for keywords associated with cross-sectional methods, then filtered to keep only studies with no longitudinal or quasi-experimental component. Designs that can carry causal weight under explicit assumptions, such as instrumental variables or regression discontinuity, were deliberately excluded. What remains is the set of papers with the least room to argue.
That 46.3 percent is a classifier output, not 194,631 human readings. A fine-tuned BERT model did the coding, trained on GPT-4o labels that were themselves validated against expert coders, and the validation points one way rather than both. Against human coders the GPT labelling ran high on precision and lower on recall, which the authors read as a slight underestimate of how much causal language is actually there. When two of them manually reviewed a sample from prestigious venues, 68 percent met strictly cross-sectional criteria, so that subset carries some misclassification in the other direction. The descriptive analysis was not preregistered.
Within the corpus, the pattern is not confined to obscure outlets. Among the 127,598 articles the team could link to SCImago Journal Rank data, causal language was more common higher up: 54.4 percent of abstracts in journals above the median rank of 3.03, against 43.3 percent below it. Twenty elite journals across four disciplines came in at 43.2 percent, which the paper calls slightly lower than the overall rate. By 2024 business research led at 84 percent, with economics at 67, psychology at 55, political science at 53 and sociology at 46. A newer journal that explicitly disallows causal claims, the Journal of Quantitative Description: Digital Media, still carried them in 12.9 percent of its articles.
Going past abstracts complicates the picture rather than settling it. Across full-text snippets from 69,580 articles — the 35.7 percent of the corpus whose full text was available, which the authors note is not a randomly distributed subset — 81.8 percent contained at least one causal snippet, well above the abstract rate. But when the team matched individual claims to their abstract counterparts across 100 articles, most matched claims were associational or descriptive, and 26.8 percent were causal overreach at the claim level. The authors explain the gap themselves: full texts hold many more claims, so any one causal statement is diluted by the associational ones around it. What survives both readings is the finding that abstracts are a poor guide to the paper behind them. That hundred was built as 25 articles per abstract category, and inside it all 25 whose title or abstract carried an unhedged causal claim also carried a matching one somewhere in the full text, as did 11 of the 25 whose abstracts stayed purely descriptive.
What happened when 1,105 people were handed the abstracts
The second half of the paper is where the reader arrives. The team preregistered an experiment, recruited 1,105 US adults holding at least a bachelor’s degree through Prolific, and gave each of them one of 28 real abstracts from top-tier journals in one of four versions: the original with its causal phrasing (216 people); a rewrite carrying the same content in strictly associational language (240); the original plus a short methodological note stating the study was cross-sectional and could not on its own establish causality (229); or that note plus labelled “AI” feedback highlighting the interpretive problems (202). A fifth, exploratory group of 218 answered the belief questions before reading an abstract, then read one and answered them again.
Asked directly whether the study provided causal evidence, participants moved. All three interventions cut agreement, with the methodological label at β = −0.4 and associational wording at β = −0.3, effects the authors describe as medium-sized, all at P ≤ 0.003 against a threshold of 0.05. On the paper’s primary measure, telling people moved the needle without turning it around: agreement fell in every intervention arm but stayed on the agree side of the seven-point scale.
Then each participant was asked to summarize, in their own words, what the study showed, and the summaries went the other way. Those free-text answers were themselves machine-coded for causal phrasing, and in the original condition 71.3 percent came back causal. With the note, 64.2 percent still did, a reduction that did not reach significance on the researchers’ own preregistered test. Two conditions did produce significant reductions: the “AI” feedback at 59.4 percent, and the rewrite, which removed the causal wording from the abstract rather than arguing with it, at 49.6 percent. Half the participants wrote a causal summary anyway, of a text that had made no causal claim at all.
A third measure went further still. Asked indirectly, about whether changing the independent variable would produce a change in the outcome, none of the three interventions produced evidence of a significant difference. Against the group who reported beliefs before reading anything, the abstract shifted those beliefs only slightly.
The sample is worth holding onto. These were college-educated Americans recruited through a research platform and, so far as the authors know, no practising social scientists; age and sex were not collected.
Instruction worked on the models
Run through GPT-4.1, the same stimuli produced causal claims in 99.3 percent of summaries of the original abstracts. Unlike the humans, the model moved a long way under instruction: 89.8 percent under the rewrite, 41.7 percent with the methodological note, 30.3 percent with the “AI” feedback. None of it went to zero, and the authors say so plainly: no intervention fully eliminated causal framing.
A follow-up widened this to five models, GPT-4.1, GPT-5.1, GPT-5.4, DeepSeek Chat-V3.2 and Anthropic’s Claude Sonnet 4.6, across four summarization prompts, on 100 articles each. Prompts asking for plain language at an eighth-grade reading level, or for real-world significance, raised the rate of unhedged causal claims above the rate in the abstracts themselves, across every model. A prompt explicitly asking for caution reduced causal language in every model except GPT-4.1, and averaged 5.0 percent across the set. (The paper’s own figure caption reports that reduction across models with no exception noted, which its main text does not; the main-text version is used here.) Stuart-Maxwell tests found no evidence of a difference between summaries built from full texts and summaries built from abstracts alone, for any prompt or model combination.
These summary codings rest on classifiers built for the job rather than on the corpus model: the human summaries were coded with an updated version of the causal-language prompt, and the five-model follow-up used a new classifier tailored to summaries, whose validation sits in the Supplementary Information that was not fetched for this piece.
The two experiments answer different questions, which the authors are careful to say, so none of this ranks machines against people.
What it does do is order the interventions. Removing the causal wording before the reader saw it worked. Attaching a note and trusting an adult reader with it moved the judgement they were asked to make, and not the sentence they then wrote. Asked to be careful, the model was careful; the people went on reading the sentence they expected to find.
The practical upshot is not a rule for spotting bad science, and it is certainly not a reason to disbelieve an observational result — the estrogen story runs both ways, and a patient weighing hormone therapy should be having that conversation with her doctor. Every limit above is the paper’s own, lifted from its Discussion and Methods; nobody here re-ran a classifier or re-coded an abstract, and the Supplementary Information, where several of the robustness checks live, was not fetched for this piece. The preregistration and the analysis code and stimuli are both public, which is more than most of the 194,631 can say.
A causal reading is what the mind reaches for first, and an associational one has to be held in place by effort. Most of the fixable part of this therefore sits with the person writing the sentence.