Correction, 15 September 2026: The headline and opening previously treated the absence of the phrase ‘human in the loop’ as the absence of a legal human-oversight mandate. Article 14 expressly requires human oversight for high-risk AI systems, while leaving the measures proportionate to risk, autonomy and context. The article has also removed an unverified editorial note.

The EU AI Act does not use the expressions ‘human in the loop’ or ‘human on the loop’. That vocabulary is common in procurement documents, vendor material and compliance presentations, but the wording is not the legal test.

Article 14 of Regulation (EU) 2024/1689 is expressly titled ‘Human oversight’. It requires high-risk AI systems to be designed and developed so that natural persons can effectively oversee them while they are in use. The measures must be proportionate to the risks, the system’s autonomy and its context. Overseers must be able to understand the system’s capacities and limits, guard against automation bias, interpret its output, disregard or override it, and interrupt the system. For specified biometric-identification decisions, separate verification by at least two competent natural persons is generally required.

Article 26 also requires deployers of high-risk systems to assign human oversight to people with the necessary competence, training and authority. The Act therefore does mandate human oversight. What it does not impose is one universal architecture in which a person must approve every output before a system can act. Whether oversight is sufficient depends on the system, risk, instructions and use context.

Where the phrase actually comes from

In November 2012, Human Rights Watch and Harvard’s International Human Rights Clinic published a report arguing for a pre-emptive ban on fully autonomous weapons. It is called Losing Humanity, and the vocabulary now used to reassure hospital boards was written in it.

Human in the loop: a machine that selects targets and delivers force only on a human command. Human on the loop: a machine that selects and delivers under the oversight of a human who can override it. Human out of the loop: a machine that does both with no human involvement.

Human Rights Watch codified and popularised that taxonomy rather than inventing it, since in and out of the loop are older control-systems terms. What the report did was attach moral weight to the preposition. And it included a warning that has aged uncomfortably well: an on-the-loop system with only nominal supervision is, in practice, effectively autonomous.

Fourteen years later, two of those three phrases are used interchangeably in AI marketing, by people who would be startled to learn they once described the conditions under which a machine may kill someone.

What changes between in and on

In the loop, the default is inaction. Nothing happens until a person decides it should. The human is a gate, and the system’s throughput is limited by how fast that person can think.

On the loop, the default is action. The system proceeds, and the person intervenes if something is wrong. The human is a brake, and the system’s throughput is unlimited.

Being in the loop means the system waits for you. Being on the loop means the system has already gone, and you are the reason it might come back.

That difference is the entire commercial appeal of the second arrangement, and it is also why the oversight it provides degrades so reliably. A gate has to be opened, and the opening leaves a record. A brake only has to be available, and availability is very difficult to audit. Nominal supervision and real supervision look identical on an org chart.

What the Act leaves flexible

The Act regulates outcomes and capabilities rather than choosing one loop model for every high-risk system. Article 14 makes oversight measures proportionate to the risks, autonomy and context of use. That flexibility can accommodate pre-approval, exception review, stop controls and other safeguards, but it does not make any one staffing arrangement automatically compliant. Providers must design the measures and identify them in the instructions for use; deployers must assign qualified people and monitor operation in accordance with those instructions.

The US Department of Defense, dealing with the same problem in the domain the vocabulary came from, went further and refused the loop framing outright. Directive 3000.09 requires appropriate levels of human judgment over the use of force. That is either an honest admission that the binary was always too crude, or a way of never having to say which side of it you are on.

The keynote where this stopped being abstract

The reason I have been reading a fourteen-year-old weapons report is a keynote at this year’s European Health Psychology Society conference, given by Alex Gillespie of the London School of Economics, titled “Using AI to analyze patient voice and hospital listening: insights into patient safety.”

The underlying work is nearly a decade old and involves no AI at all. Gillespie and Tom Reader built the Healthcare Complaints Analysis Tool, published in BMJ Quality and Safety in 2016, to code what patients and families actually say when they complain. Their later work across 1,110 complaints from 56 NHS trusts showed that complaints surface problems incident reporting misses, and a subsequent analysis across 59 trusts found that the severity of clinical problems described in complaints was the only patient-generated measure associated with hospital-level mortality.

That association is at hospital level and cross-sectional — one step short of proving complaints predict deaths. The real finding is a safety signal sitting in complaint letters, with nobody around to read it at scale.

Which is where a language model becomes the obvious answer, and where the preposition stops being a vocabulary question.

The finding that should slow everybody down

The strongest evidence here comes from Gillespie’s own group, which is what makes it worth taking seriously. A team including Hannah Bunt, Alex Goddard, Reader and Gillespie validated GPT-4o against human coders on three classification tasks of 1,500 items each, drawn from NHS complaint data

On identifying reported speech, whether a passage is quoting something someone said, the model achieved an F1 score of 0.91. On identifying repair, the same. Near-human agreement.

On classifying harm, whether the complaint describes a patient having been harmed, weighted kappa ranged from 0.49 to 0.57 depending on the prompt, and the model didn’t just disagree with human coders — it systematically overestimated harm. Moderate at best. Unremarkable in exploratory research, and not what anyone wants underneath a triage decision.

The pattern is the point. The model is reliable where the language is explicit and unreliable where the judgement is clinical. It is best at the parts that matter least and worst at the part the whole exercise exists for. Nobody should write that AI reads patient complaints as well as humans do, because the paper says close to the opposite precisely where it counts.

That is the finding that makes the oversight design matter. A system whose weakest classification is harm may need a person to act as a gate rather than merely a brake. The Act does not decide that architecture in the abstract; its proportionality test requires the safeguards to fit the risk and context.

The second risk, which nobody regulates

The other thing raised at that keynote had nothing to do with oversight architecture, and it is the one I have not been able to put down.

The risk is that a research culture rewarding publications rather than outcomes will validate these systems on the metrics that publish. An F1 of 0.91 on reported-speech classification is a publishable number. A weighted kappa of 0.57 on harm is a footnote. Both are in the same paper, which is to the authors’ credit, and only one of them is going to survive into the procurement deck.

This is the paper-versus-impact problem in its most concrete form: which of two honest numbers gets carried forward, decided by an incentive structure that has no view on patient safety at all, with no fraud or hype required.

Where the preposition ends up

So it is worth asking directly, of any system that advertises human oversight: does the system wait, or does it go? If it goes, who is watching, how many outputs an hour, and what happens to that person if they intervene too often?

Article 14 does not answer those questions with one universal workflow. Providers and deployers still have to make the oversight effective, proportionate and usable by people with the required competence, training and authority.

What stays with me is that a distinction developed to govern the use of lethal force now silently governs how a complaint letter from a bereaved family gets triaged, that the 2012 warning about nominal supervision travelled with it word for word, and that the phrase everybody quotes to prove a person is still in charge turns out not to appear in the law they are quoting.