Writing / AI and Automation
Human-in-the-Loop Is Not a Checkbox
Ben Tartaglia · Published October 2, 2026
“A human reviews it” sounds reassuring. I want to know what that human can actually see, decide, and stop.
A review step can be useful, or it can become a decorative pause before the system does what it was going to do anyway. The difference lies in the reviewer’s information, authority, and available time.
Imagine a hypothetical AI-generated recommendation to change a customer’s access permissions. A screen shows the proposed action and a green approval button. If the reviewer cannot see the original request, the relevant access policy, and the uncertainty behind the recommendation, they’re being asked to approve a conclusion without its evidence.
Meaningful review would put that information within reach. It would explain why the case was routed to a person, identify the action’s consequence, and allow the reviewer to reject it, request more information, or escalate it. The resulting decision should leave a record someone else can follow.
I want the gate placed before the consequential action. Reviewing a message after it has been sent may help us learn, but it cannot prevent that particular mistake. Some tasks justify sampling after completion. Others need a decision before execution. The consequence should determine the placement.
The Glowing Oracle Lab gives this principle a concrete test in my own work. Its current records include two evaluations marked as requiring human review, and the backend check found no recorded human reviews. That is an open review obligation. The presence of the flag doesn’t mean oversight has happened.
It also tells us something about how a useful dashboard should behave. Review-required items need a visible queue, a named owner, and a defined disposition. Otherwise the flag is a field in a database waiting patiently for someone to remember it exists.
Automated evaluation can help focus attention. A second model can compare outputs against a rubric and identify disagreements. That adds another assessment; it doesn’t transfer accountability away from the people operating the system. The evaluator’s criteria and behavior need examination too.
Review design has a capacity problem as well. If every trivial item requires approval, the queue becomes noise. If hundreds of consequential decisions are compressed into a few minutes, the organization has created pressure to rubber-stamp them. I’d route review according to consequence, uncertainty, novelty, and known failure patterns, then measure whether the queue can actually be worked.
Useful measures include review age, rejection reasons, escalation frequency, and the errors found through sampling. Approval rate alone tells us very little. A high rate could mean excellent output, weak review, or a team without enough time to investigate.
The human should be able to change the outcome. They should have evidence, enough time, and a path to stop or escalate the work. The system should preserve what they decided and why.
That’s the standard I mean when I say human judgment belongs in an AI workflow. It is a responsibility we design for and complete.