Everybody attributes the phrase to the Act. Procurement documents cite it. Vendors promise it. Compliance decks put it on a slide with the article number next to it.
Article 14 does not contain it. Not once, and neither does it contain ‘human on the loop’.
What Article 14 requires instead is a list of capabilities in whoever does the overseeing. They must properly understand the system’s capacities and limits. They must remain aware of automation bias, the documented tendency to over-trust automated output. They must be able to correctly interpret what the system produces, to decide not to use it, to disregard or override its output, and to interrupt it, including through a stop button. For certain biometric identification uses, two separate natural persons must confirm.
Read that list again and notice what is missing from it. Every item describes a competent person. Not one of them describes when the system is allowed to act. A person can satisfy every requirement on that list while reviewing exceptions after the fact, and the Act does not say otherwise.
Where the phrase actually comes from
In November 2012, Human Rights Watch and Harvard’s International Human Rights Clinic published a report arguing for a pre-emptive ban on fully autonomous weapons. It is called Losing Humanity, and the vocabulary now used to reassure hospital boards was written in it.
Human in the loop: a machine that selects targets and delivers force only on a human command. Human on the loop: a machine that selects and delivers under the oversight of a human who can override it. Human out of the loop: a machine that does both with no human involvement.
Human Rights Watch codified and popularised that taxonomy rather than inventing it, since in and out of the loop are older control-systems terms. What the report did was attach moral weight to the preposition. And it included a warning that has aged uncomfortably well: an on-the-loop system with only nominal supervision is, in practice, effectively autonomous.
Fourteen years later, two of those three phrases are used interchangeably in AI marketing, by people who would be startled to learn they once described the conditions under which a machine may kill someone.
What changes between in and on
In the loop, the default is inaction. Nothing happens until a person decides it should. The human is a gate, and the system’s throughput is limited by how fast that person can think.
On the loop, the default is action. The system proceeds, and the person intervenes if something is wrong. The human is a brake, and the system’s throughput is unlimited.
Being in the loop means the system waits for you. Being on the loop means the system has already gone, and you are the reason it might come back.
That difference is the entire commercial appeal of the second arrangement, and it is also why the oversight it provides degrades so reliably. A gate has to be opened, and the opening leaves a record. A brake only has to be available, and availability is very difficult to audit. Nominal supervision and real supervision look identical on an org chart.
Why the omission is probably deliberate
I do not think the drafters forgot. Requiring a human approval step on every high-risk output would make many of these systems useless, and an unused system attracts no regulation at all — it just gets shelved. Article 14’s requirements are explicitly proportionate to the risk and to the autonomy of the system, which is a sensible piece of engineering and also the exact mechanism by which an organisation can stay on the loop and still tick every box.
The US Department of Defense, dealing with the same problem in the domain the vocabulary came from, went further and refused the loop framing outright. Directive 3000.09 requires appropriate levels of human judgment over the use of force. That is either an honest admission that the binary was always too crude, or a way of never having to say which side of it you are on.
The keynote where this stopped being abstract
The reason I have been reading a fourteen-year-old weapons report is a keynote at this year’s European Health Psychology Society conference, given by Alex Gillespie of the London School of Economics, titled “Using AI to analyze patient voice and hospital listening: insights into patient safety.”
The underlying work is nearly a decade old and involves no AI at all. Gillespie and Tom Reader built the Healthcare Complaints Analysis Tool, published in BMJ Quality and Safety in 2016, to code what patients and families actually say when they complain. Their later work across 1,110 complaints from 56 NHS trusts showed that complaints surface problems incident reporting misses, and a subsequent analysis across 59 trusts found that the severity of clinical problems described in complaints was the only patient-generated measure associated with hospital-level mortality.
That association is at hospital level and cross-sectional — one step short of proving complaints predict deaths. The real finding is a safety signal sitting in complaint letters, with nobody around to read it at scale.
Which is where a language model becomes the obvious answer, and where the preposition stops being a vocabulary question.
The finding that should slow everybody down
The strongest evidence here comes from Gillespie’s own group, which is what makes it worth taking seriously. A team including Hannah Bunt, Alex Goddard, Reader and Gillespie validated GPT-4o against human coders on three classification tasks of 1,500 items each, drawn from NHS complaint data
On identifying reported speech, whether a passage is quoting something someone said, the model achieved an F1 score of 0.91. On identifying repair, the same. Near-human agreement.
On classifying harm, whether the complaint describes a patient having been harmed, weighted kappa ranged from 0.49 to 0.57 depending on the prompt, and the model didn’t just disagree with human coders — it systematically overestimated harm. Moderate at best. Unremarkable in exploratory research, and not what anyone wants underneath a triage decision.
An independent replication points the same way, with substantial agreement on which domain a complaint belongs to and poor agreement on severity and harm. I would verify that citation before leaning on it, but it is consistent with the validation study.
The pattern is the point. The model is reliable where the language is explicit and unreliable where the judgement is clinical. It is best at the parts that matter least and worst at the part the whole exercise exists for. Nobody should write that AI reads patient complaints as well as humans do, because the paper says close to the opposite precisely where it counts.
And that is the finding that makes the missing phrase matter. A system whose weakest classification is harm is a system whose human should be a gate. What the Act permits is a brake.
The second risk, which nobody regulates
The other thing raised at that keynote had nothing to do with oversight architecture, and it is the one I have not been able to put down.
The risk is that a research culture rewarding publications rather than outcomes will validate these systems on the metrics that publish. An F1 of 0.91 on reported-speech classification is a publishable number. A weighted kappa of 0.57 on harm is a footnote. Both are in the same paper, which is to the authors’ credit, and only one of them is going to survive into the procurement deck.
This is the paper-versus-impact problem in its most concrete form: which of two honest numbers gets carried forward, decided by an incentive structure that has no view on patient safety at all, with no fraud or hype required.
Where the preposition ends up
So it is worth asking directly, of any system that advertises human oversight: does the system wait, or does it go? If it goes, who is watching, how many outputs an hour, and what happens to that person if they intervene too often?
Article 14 will not ask those questions for you. It describes a competent overseer and leaves the posture to whoever is buying.
What stays with me is that a distinction developed to govern the use of lethal force now silently governs how a complaint letter from a bereaved family gets triaged, that the 2012 warning about nominal supervision travelled with it word for word, and that the phrase everybody quotes to prove a person is still in charge turns out not to appear in the law they are quoting.