Oversight has a capacity
Every AI governance plan assumes someone is watching. Almost none of them ask how much watching a person can actually do.
Every AI governance plan assumes someone is watching. Almost none of them ask how much watching a person can actually do.
If your AI governance rests on a human in the loop, you have made a claim about human capacity, and you have almost certainly never checked it. That claim is testable. Most organizations would fail the test today.
I look at this the way I look at any job: what is the work, who is doing it, how is the role designed, and does the organization around it permit the behavior we are counting on. Those are industrial-organizational questions, and they get asked about warehouse throughput and call center quality all the time. They rarely get asked about the reviewer whose signature is the entire control.
Governance assumes uniform attention. People do not supply it.
The reviewer in your risk register is an average person having an average day. No such person exists.
Attention to rare events degrades within the first half hour of a monitoring task. Norman Mackworth established that in 1948 with radar operators, and it has been replicated for seventy years since. Layer on what Parasuraman and Riley documented about automation: when a system is usually right, people stop scrutinizing it. Your model's accuracy is what erodes the review.
Then add the variance that any I/O practitioner expects and no governance document accounts for. Two reviewers with the same title differ substantially in how much they catch. The same reviewer differs from morning to late afternoon, and again under time pressure, and again when the queue is visible and growing. Expertise helps with judgment and does very little for sustained vigilance. Conscientiousness predicts effort, and effort is precisely what the decrement defeats.
Governance written for the average of those people is written for nobody. It will be right on paper and wrong in the room.
Org design decides whether oversight is possible at all
Even a well-rested, well-matched reviewer cannot exercise oversight the structure forbids.
Watch what happens when ten agents are pointed at one approver. On the org chart the function is covered. In the work itself the reviewer shifts, without ever choosing to, from judging to glancing, because the volume leaves no other option. Plausible output passes. The rare genuine error arrives looking exactly like everything else, and it arrives after attention has already worn thin.
Three structural conditions decide whether that reviewer can act on what they see. Whether stopping the line costs them anything. Whether escalation goes to someone who can act, or into a queue. Whether the metric they are judged on is throughput, which every reviewer learns within a week.
Role ambiguity does the rest. When several people are nominally accountable for reviewing the same output, each one reasonably assumes the others are looking closely. That is diffusion of responsibility, it is one of the oldest findings in social psychology, and RACI exists specifically to prevent it. Two accountable parties produce zero.
None of that is a technology problem, which is why more tooling rarely fixes it.
Measure the load, then design the role to fit it
HAIL-ETHIC treats oversight as a designed human system with a measurable load. Four things follow, and you can start all four this week.
State the ratio. Decisions per reviewer per shift, written down. If nobody in the organization can produce that number, you are running on an assumption rather than a control.
Set a ceiling and defend it. A ratio that looks responsible on a slide becomes indefensible when volume triples, and volume always triples. Decide the ceiling before the growth, because you will not decide it during.
Sample with depth instead of reviewing everything shallowly. Exhaustive glancing catches less than structured sampling with real attention on the sampled cases.
Escalate on load, not only on outcome. Most thresholds fire after something goes wrong. Add one that fires when the queue exceeds the ceiling, because that is when the control has quietly stopped working, and it is visible well before the incident.
Then design the job so a competent person can hold it. Rotation off the monitoring task before the decrement bites. Hard load limits with teeth. Decision support that surfaces the cases most worth a second look. Real authority to stop the line, exercised at least once without career consequence, so everyone learns it is genuine. Oversight that depends on heroics is a postmortem waiting for a date.
The question worth taking to your next governance review
Stop asking whether there is a human in the loop. Ask what that human's capacity is, how much it varies across your actual people, and whether your structure lets them use it.
The first question an org chart answers. The second requires you to know your volumes, your role design, and your escalation paths. It will tell you something uncomfortable, which is the reason it is worth asking.
If you cannot answer it today, you have your finding. Start there.
Sources and further reading
- Mackworth, N. H. (1948). The breakdown of vigilance during prolonged visual search. Quarterly Journal of Experimental Psychology, 1(1), 6-21.
- Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253.
- Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381-410.
- NIST AI Risk Management Framework, MEASURE function, on evaluating the effectiveness of human oversight.