Anyone who has driven a car with enough driver assistance to make it feel almost self-driving knows the feeling. The car holds the lane, keeps its distance, and mile after uneventful mile, your attention begins to loosen. You are still responsible for what happens next. If something went wrong in the next three seconds, you would be the human in the loop, expected to take over just when you had the least reason to be alert.
I have been thinking about that seat because the same arrangement is showing up everywhere else. For years the standard reassurance about artificial intelligence has been to keep a human in the loop. Let the software do the analysis or the drafting, as long as a person signs off before anything important happens. The newer tools, though, do more than answer questions. Your email program drafts the reply and waits for you to hit send. At work, AI tools can sort applicants, analyze information, and write up recommendations, leaving the person's part increasingly concentrated at the end: reviewing the work and clicking approve. On paper, the human never left the loop, but being in the loop and exercising judgment in it are not the same thing.
A watchman who is only present
Ezekiel describes a watchman set on the city wall with a simple responsibility: if he sees the sword coming, he sounds the warning; if he sees it and says nothing, the consequences are on him (Ezekiel 33:2–6). His value is not that he occupies the wall, but that he is watching closely enough to recognize danger and respond.
That is what makes the image so relevant to AI oversight. Once a person is placed at the end of a process, everyone involved can feel a little safer. The company selling the software can say a human reviews every result, the office using it can say no decision is fully automated, and the letter that lands in your mailbox can assure you that your case was reviewed. But the person at the desk inherits all of that confidence along with a file that already looks finished. A sloppy file naturally invites questions; one with tidy numbers and a reasonable explanation is much easier to sign, especially when questioning it requires the very effort the software was meant to save.
Watching is harder than it looks
People who study this for a living noticed the problem long before anyone had a chatbot. In 1983 the psychologist Lisanne Bainbridge described what she called the ironies of automation. As a machine takes over more of a task, the person is left with the job of watching for the rare failure, and watching for something that almost never happens is a job people do badly (Bainbridge, 1983). That is the driver on the highway again.
A review of 106 experiments found something surprising: combining a person with AI did not always produce the best result. People using AI generally performed better than people working alone, but the combination often performed worse than the AI could have performed by itself. In other words, simply adding human oversight did not guarantee a better decision (Vaccaro et al., 2024).
Ben Green found a similar problem in government policies. After reviewing 41 rules that require humans to oversee automated decisions, he argued that having a person involved can create more confidence than protection. A decision may technically have been “reviewed by a human” even when that person was not in a good position to catch what the system got wrong (Green, 2022).
We have put a watchman on the wall without asking whether he can see from there.
What a finished list leaves out
If you have applied for a job in the past few years, there is a good chance software played some role before a person ever saw your résumé. Now imagine sitting on the other side of the desk, where two hundred people applied and the screen presents the ten strongest candidates. The list looks impressive, but it does not show you who was ranked eleventh and why, whether the system favored applicants whose language closely matched the job posting, or whether a strong candidate was pushed down because she had taken a year away from work to care for her mother. Unless the system surfaces those details, they are easy to miss, and a list that looks complete can quietly persuade us that there is nothing left to question.
The skill this calls for is different from simply knowing how to use the tool. The recommendation may be right, and good judgment does not mean looking for reasons to reject everything AI produces. It means asking what the recommendation is leaning on, what may be missing, and whose situation never made it onto the screen. I find one question especially useful: What would have to be true for this to be wrong? And would anything in front of me have told me? That kind of scrutiny is easy to overlook because sometimes its only visible result is a decision to pause, question, or revise the first answer.
Test, then hold
Paul tells the Thessalonians to "test everything; hold fast what is good" (1 Thessalonians 5:21). The order matters. Testing comes first, and holding fast is reserved for what survives the test, which is a long way from receiving something persuasive and giving it a final glance. It is the difference between reading the lease and initialing every page where the agent points. The responsibility does not move to the machine along with the work. If anything, the more of the work the machine does, the more it matters to know which part is still ours. An AI system can produce a recommendation, but it cannot answer for choosing it, and neither can a signature at the bottom of a page.
A better question
“Was there a human in the loop?” may be turning into the wrong test. A better question, whether you run a company or just opened an email that offered to write your reply, is: Where did you actually have to think? What did you check? What could you have changed? And if you accepted the answer, did you know enough to say why?
None of this is an argument for doing by hand what a machine does well, but efficiency should free judgment rather than remove the occasions that build it, and an approve button can remove them very efficiently. There will be more AI in the loop next year than this year. What I want to know is whether the person in there with it can still see enough to judge, because nobody was ever put on a wall just to be there.
References
Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. https://doi.org/10.1016/0005-1098(83)90046-8
Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://doi.org/10.1016/j.clsr.2022.105681
Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1