The famous marshmallow test presents a child with a deceptively simple choice: eat one treat now, or wait for an adult to return and receive two. The longer a child waits, the familiar story goes, the more self-control that child possesses.

There is an assumption hidden inside the bargain. Waiting is worthwhile only if the adult is likely to keep the promise. If the second marshmallow may never arrive, taking the certain reward is not necessarily a failure of willpower. It may be the sensible choice.

A University of Rochester experiment made that assumption visible. Before children encountered the marshmallow, an adult either kept two small promises or broke both of them. Children who had just seen a reliable adult waited an average of 12 minutes. Those who had seen an unreliable adult waited about three.

The result did not show that self-control is irrelevant. It showed that the task cannot cleanly separate self-control from a child’s estimate of whether waiting will pay.

How a nursery-school task became a character test

Walter Mischel and colleagues developed the classic delay-of-gratification experiments at Stanford’s Bing Nursery School in the late 1960s and early 1970s. A preschooler faced an immediately available reward and a larger reward that required waiting alone.

The original work was more nuanced than its popular retelling. It examined how attention, distraction and the way children mentally represented a reward changed their ability to wait. Later follow-ups associated longer waiting with adolescent outcomes, and the test gradually became a cultural symbol of discipline and future success.

That symbol encouraged a tempting inference: the child who waits possesses a durable inner strength, while the child who eats early lacks it. A few minutes beside a sweet came to look like a window into character.

But the choice contains more than temptation. It contains an offer from another person. The larger reward is deferred, so the child must judge both their own ability to wait and the credibility of the adult making the promise.

The adult first established whether promises meant anything

The Rochester study by Celeste Kidd, Holly Palmeri and Richard Aslin recruited 28 children aged from three years and six months to five years and ten months. Their average age was four and a half. The researchers randomly assigned 14 to a reliable condition and 14 to an unreliable one, balancing the groups by age and gender.

Before the food task, each child completed an art project. The available crayons were old and unattractive. An experimenter offered to leave the room and fetch a better set.

In the reliable condition, the experimenter returned with a large tray of appealing art supplies. In the unreliable condition, she came back empty-handed, apologised and said she had made a mistake because there were no other art supplies.

The sequence was then repeated with stickers. The child could use one small sticker immediately or wait while the experimenter fetched a larger selection. Again, the reliable adult delivered what she had offered, while the unreliable adult returned without it.

This was not a questionnaire about trust or a comparison between children from different homes. It was a randomised manipulation in which the same basic task supplied two pieces of immediate evidence about one adult’s reliability.

Three minutes after broken promises, 12 after kept ones

Once the art materials had been cleared away, the experimenter placed a single marshmallow about ten centimetres from the table’s edge. She told the child that it could be eaten immediately, but that waiting until she returned would earn two marshmallows.

The maximum wait was 15 minutes. Children in the unreliable condition waited an average of 181.57 seconds, or three minutes and two seconds. Those in the reliable condition waited 722.43 seconds, or 12 minutes and two seconds.

The difference was not only a product of a few extreme values. Just 1 of the 14 children in the unreliable group waited the full 15 minutes. In the reliable group, 9 of 14 lasted until the experimenter returned.

A short sequence of kept or broken promises had shifted average waiting time by nine minutes. The children had not acquired or lost a stable personality trait during the art project. What changed was the evidence available for deciding whether a delayed reward was credible.

The result is sometimes summarised as proving that the marshmallow test measures trust rather than self-control. That goes too far. Waiting beside an attractive treat can still require attention control, emotional regulation and strategies for coping with temptation. Reliability changed the behaviour dramatically, but it need not be the only influence on it.

A rational choice still requires self-regulation

Imagine the decision from the child’s position. One marshmallow is certain and available. Two marshmallows are more valuable, but their expected value depends on the probability that the adult will return with the promised reward.

When the adult has just delivered better crayons and stickers, waiting looks like a good investment. When the adult has twice promised something and returned empty-handed, the same wait carries a real risk. Eating now protects the only reward the child knows exists.

This is rational in the everyday sense, not proof that four-year-olds calculate formal probabilities. Children continuously learn which people and environments are dependable. Their behaviour can reflect those learned expectations without a conscious equation.

Self-control enters after the future reward is judged worth pursuing. A child who believes the adult may still need to look away, sing, cover the treat or occupy their attention. The task therefore mixes at least two questions: “Can I wait?” and “Why should I believe waiting will work?”

That distinction echoes a broader point from Silicon Canals’ earlier look at effortless self-control. Behaviour does not emerge from willpower in isolation. Habits, cues and the structure of the situation help determine when effort is needed and whether effort appears worthwhile.

A larger extension found the effect again, but smaller

The original Rochester experiment was striking, but 28 children is a small sample and the observed effect was unusually large. In 2020, researchers at the University of Michigan repeated the reliability procedure with 60 children aged three to five.

The later study moved testing from an unfamiliar laboratory into rooms at the children’s schools. It again found evidence that children exposed to a reliable experimenter waited longer than those exposed to an unreliable one. The effect was smaller than in the Rochester experiment, and the analysis found an interaction between gender and condition.

That is a healthier result than either total failure or a perfectly identical replication. It supports the central claim that evidence about reliability can influence waiting, while warning against treating 12 versus three minutes as a fixed law of childhood.

It also narrows what the research licenses us to say. The experiment manipulated the reliability of one unfamiliar adult over a few minutes. It did not directly measure the stability of a child’s home, family income, food security or general trust in adults. Those factors may matter, but the study itself should not be used to diagnose them from one child’s choice.

The prediction-of-success story also became less tidy

The marshmallow test acquired much of its fame from the idea that early waiting forecast later life. Newer longitudinal work has found a more modest and context-dependent pattern.

A 2018 conceptual replication involving 918 children did find that waiting longer at age four was associated with better academic achievement at 15. But the raw association was about half the size reported in the earlier work. Adding controls for family background, early cognitive ability and the home environment reduced it by roughly two-thirds.

Associations with behavioural outcomes at age 15 were smaller and rarely statistically significant. The authors also found that most of the achievement difference lay between children who waited at least 20 seconds and those who did not, rather than in a smooth ladder where every extra minute signalled more future success.

A 2024 preregistered study followed 702 participants to age 26. Waiting time showed modest unadjusted relationships with educational attainment and body mass index, but almost all associations with adult achievement, health and behaviour became statistically nonsignificant after adjustment for childhood characteristics.

This does not make self-control unimportant. Nor does it erase evidence from broader, repeated measures of childhood self-control. As Silicon Canals has previously reported, a longitudinal study that assessed more than 1,000 children repeatedly between ages three and 11 found associations with later health and financial outcomes.

Those are different kinds of evidence. A broad multi-year measure of behaviour is not interchangeable with a single encounter involving one adult, one reward and one afternoon.

One marshmallow cannot diagnose a child

The Rochester result is most useful as a warning against overinterpreting behaviour. A child who eats early may be hungry, may value the second treat less, may struggle with attention, may distrust the experimenter, or may simply have learned that “later” is not always dependable.

The task records the final choice, not a pure reading of the mechanism underneath it. Labelling that child impulsive mistakes an outcome for an explanation.

There is also a practical lesson for adults, though not a recipe for turning every promise into a self-control intervention. Delayed rewards become easier to choose when the world has demonstrated that they arrive. Consistency makes patience less speculative.

That reframes the marshmallow test without discarding it. The child is still regulating attention and desire, but also making a decision under uncertainty. Waiting can show self-control. Eating can show that the offer did not look trustworthy enough to justify the cost.

A second marshmallow is valuable only if it appears. Before we judge a child’s willpower, the Rochester experiment asks us to inspect the promise.