When Human-in-the-Loop Becomes the Bottleneck
Human review does not scale automatically with AI output. At sufficient volume, queues, alert fatigue, shallow checking, and rubber-stamping can turn a useful safeguard into the constraint—or into a ceremonial step. The response is not to remove people wholesale, but to distinguish genuinely new judgment from repeatable intervention and design evidence for each.
By Ken Ohyama, Founder · Published September 9, 2026 · Reviewed September 9, 2026
- human-in-the-loop
- review queues
- AI operations
The machine can add work while the reviewer is still in the meeting
A coding assistant proposes another pull request. A security tool produces another alert. A support model drafts another reply. The practical constraint is rarely whether the system can produce one more item. It is whether the people trusted to notice strange and consequential cases can examine one more item with real attention.
A queue is not just an inconvenience. It records the difference between the rate at which work is created and the rate at which judgment can be applied.
Different domains show different versions of the same pressure
METR reported in March 2026 that roughly half of AI-authored pull requests that passed the SWE-bench automated grader still would not have been merged by actual maintainers. The automated grader was 24.2 percentage points more optimistic than maintainers. That result is about this benchmark setting, not every coding agent. It is nevertheless a useful reminder that passing an automated check and satisfying a person responsible for a codebase are different tests.[Many SWE-bench-passing PRs would not be merged into main]
Project Aurora describes estimates that 90–95% of transaction-monitoring alerts can be false positives. Microsoft cited research finding 46% of security alerts were false positives and 42% went uninvestigated. These are distinct domains and methods; together, they show why volume can overwhelm a review layer before anyone has decided how to learn from it.[Project Aurora][Unify now or pay later: New research exposes the operational cost of a fragmented SOC]
The capacity gap
Output can compound before review capacity changes
This is an editorial model, not a numerical forecast.
- 01
Machine output rises
Each new workflow can create another stream of drafts, alerts, and proposed actions.
- 02
Expert hours stay finite
The same experienced people still need time to understand unusual cases.
- 03
The queue accumulates
Backlog, shallow review, or unreviewed work appears somewhere in the system.
- 04
The organization chooses
Keep paying for repeated checks, weaken review, or learn which classes can safely stop returning.
A larger queue is not proof that review should disappear. It is a reason to ask what the queue is teaching.
A queue changes human behavior
When items arrive faster than they can be considered, reviewers adapt. They sample. They look for obvious signals. They approve familiar shapes. They postpone difficult cases. None of this means reviewers are careless. It means the workflow has asked finite attention to perform unbounded work.
The danger is subtler than backlog. A formal human-in-the-loop step can remain in the diagram even after it has lost its ability to discriminate. The person is still there; the decision has become increasingly pre-shaped by volume, interface, and time pressure.
The bottleneck is often repeated judgment, not all judgment
Some queues are full of cases that deserve a person every time. Others contain a growing seam of interventions that have already been made, explained, and made again. The first asks for staffing, policy, or a different operating model. The second asks whether the organization can preserve a lesson instead of repeatedly paying to rediscover it.
That distinction leads to the practical work in How to Reduce Human Review Without Removing Human Oversight.
Where Never Twice may fit
Never Twice is for organizations with a meaningful AI review queue and enough reviewed work to examine recurring intervention. It works beside the existing workflow first, using replay and shadow operation to establish what may safely leave routine human review. It is not legal advice, AI certification, a promise of autonomous operation, or a substitute for required human decisions.
Explore Never TwiceGlossary
- Alert fatigue
- Reduced ability to notice and act on important signals after repeated exposure to a high volume of alerts.
- Rubber-stamping
- Approval that occurs without the meaningful consideration the approval is meant to represent.
- Review queue
- Work awaiting human approval, correction, rejection, or escalation.
Sources & further reading
- Many SWE-bench-passing PRs would not be merged into main (opens in a new tab) · METR · 2026
- Project Aurora (opens in a new tab) · Bank for International Settlements · 2025
- Unify now or pay later: New research exposes the operational cost of a fragmented SOC (opens in a new tab) · Microsoft Security · 2026
- What is human-in-the-loop? (opens in a new tab) · IBM Think
This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.
Continue the research
What took decades to learn
should not disappear in a day.
The road ahead should remember how the company came this far.
