Insights & Resources

Before You Automate Expert Work, Find the Exceptions

Before automating expert work, leaders should map the cases in which the documented process stops being sufficient: unusual conditions, weak signals, conflicting goals, missing evidence, and thresholds that require escalation. That exception map should become test material for the proposed system. It will not prove the AI is safe or complete, but it can reveal whether the organization is automating a real decision terrain or only its routine path.

By Ken Ohyama, Founder · Published August 23, 2026 · Reviewed August 23, 2026

  • AI delegation
  • expert work
  • decision boundaries

The procedure is usually written for the middle of the road

Most process documents describe the path that can be repeated. They show the required fields, the sequence of approvals, the normal handoff, and the answer that is right often enough to become policy. That is exactly why they are useful—and why they can be misleading as a complete specification of expert work.

Experienced operators spend much of their value at the edges. They notice when two ordinary signals conflict, when a familiar customer is behaving out of character, when the input is technically valid but contextually wrong, or when following the procedure would satisfy the letter of the rule while creating a larger risk.

If an AI system is trained, prompted, or evaluated only on the documented path, the organization may learn that it can reproduce routine execution. It has not yet learned whether the system—or the people overseeing it—can recognize when routine execution should stop.

Current AI evidence makes exception work a serious design question

The International AI Safety Report 2026 describes current general-purpose AI capability as jagged across tasks and contexts. It reports continuing difficulty with long-term planning, many-step tasks, and unexpected obstacles, while also emphasizing rapid improvement and uncertainty about how capabilities will develop. Benchmarks can overstate practical utility when controlled tests do not represent dynamic real-world conditions.

That evidence does not establish that every AI system fails on exceptions. It does not show that undocumented expert knowledge is a leading cause of failed enterprise deployments. The defensible inference is narrower: where an expert workflow contains consequential unusual cases, leaders should test those cases directly rather than assume success on routine examples will generalize.

This is especially important when an AI agent can act, not merely recommend. Permissions, reversibility, monitoring, and recourse determine how far one missed exception can travel before a human sees it.

Authoritative synthesis

The International AI Safety Report 2026 describes an evaluation gap: performance in controlled pre-deployment tests does not reliably predict practical utility or risk in dynamic real-world settings, and current agents can struggle with long tasks and unexpected obstacles.

International AI Safety Report 2026 · International AI Safety Report

Method note: The report synthesizes evidence on general-purpose AI available before December 2025. Capabilities are changing quickly and vary by system, task, tools, and deployment context.

An exception is more than a rare event

Some exceptions are statistical outliers. Others occur regularly but remain difficult because the relevant facts arrive late, goals conflict, or local history changes the meaning of a signal. A case can be ordinary in frequency and exceptional in judgment.

Useful categories include boundary cases where a small fact changes the call; missing-information cases where proceeding would be premature; conflicting-authority cases where two legitimate policies point in different directions; and escalation cases where consequence or irreversibility exceeds the operator’s mandate.

The point of a taxonomy is not to create a universal list. It helps leaders sample the terrain deliberately. Without it, an evaluation set tends to become a collection of clean examples that the organization already knows how to solve.

Illustrative decision terrain

Documented process and actual work answer different questions

Documented process

  • What normally happens
  • Which fields and approvals are required
  • How a clean case moves from start to finish
  • Who owns the routine handoff

Actual decision terrain

  • Which weak signal changes the interpretation
  • When valid inputs should still be challenged
  • Which conflicting goals require judgment
  • Where consequence, uncertainty, or authority demands escalation

The procedure remains necessary. The exception map shows where following it is no longer enough.

Experts reveal the terrain through cases, not slogans

The Critical Decision Method was developed to elicit critical cues, discriminations, and judgments from experienced people working in demanding naturalistic environments. Applied Cognitive Task Analysis similarly helps practitioners identify difficult cognitive demands, cues, strategies, and potential errors, then represent them for training and design.

For automation planning, those traditions suggest a practical line of inquiry. Begin with nonroutine incidents. Reconstruct the changing situation. Ask what the expert noticed, what made the ordinary response inadequate, what evidence would have reversed the call, and when the matter had to move to someone with different authority.

This should not be treated as mind extraction. Experts can misremember, disagree, and rationalize. The cases become stronger when compared with records, near misses, failures, and the accounts of adjacent roles. Preserving disagreement may be more honest than forcing a single golden rule.

Turn the exception library into an evaluation, not a decoration

An exception register has little value if it sits beside the deployment plan as background reading. Each consequential case should become a test: Does the system detect that the ordinary path no longer applies? Does it ask for the missing information? Does it escalate to the right owner? Can it refuse an action when the consequence is high and the evidence weak?

Use contrasting pairs where two cases look nearly identical but demand different responses. Hold some cases out of development. Vary details that should matter and details that should not. Test the full system—including retrieval, tools, permissions, and human handoffs—because a capable model inside a poorly designed workflow can still produce a fragile operation.

Passing the set does not prove general reliability. It supplies bounded evidence and a clearer failure vocabulary. Production monitoring should then look for new exception classes, changes in context, and recurring escalations that indicate the map is drifting.

Pre-deployment field test

Follow the case until a human call becomes necessary

The sequence turns an expert’s exception into something a team can test across the full workflow.

  1. Routine

    Establish the ordinary path

    Show what the process expects when inputs, goals, and authority are aligned.

  2. Signal

    Introduce a weak contradiction

    Add the detail an experienced person would notice even though the case still appears valid.

  3. Exception

    Cross the decision boundary

    Change the condition that makes the routine response inadequate or unsafe.

  4. Escalation

    Test the handoff

    Observe whether the system asks, pauses, refuses, or routes the case with enough context.

  5. Human call

    Confirm authority and recourse

    Verify that the reviewer has the competence, time, information, and right to intervene.

A good evaluation asks whether the system knows where its path ends—not only whether it can complete the path.

The model may infer the exception; make it show you

A reasonable objection is that modern models can infer patterns no expert explicitly articulated. That may be true for a given task, and explicit rules can even make performance worse when they freeze an outdated interpretation. The answer is not to presume that elicitation must outperform the model. It is to run the comparison.

Test the proposed system on representative and adversarial cases with and without the elicited material. Ask subject-matter experts to examine errors as well as agreements. If the model handles a class of exceptions reliably without additional representation, the organization has learned something valuable. If it fails, the team has found a boundary before deployment rather than after consequence.

Skagway’s practitioner hypothesis is that succession artifacts built around cases, cues, and thresholds may sometimes help create stronger automation evaluations. That proposition is technically plausible and unproven. It belongs in experiments, not in promises.

Governance begins where the routine path ends

A usable escalation design names the trigger, the human owner, the information they receive, the response time available, and their authority to pause or reverse the system. “Send to human review” is not a control design if nobody can explain what the reviewer is expected to recognize.

Review intensity should follow consequence, reversibility, observability, and recourse. Requiring a senior expert to inspect every low-risk output may erase the economic value of automation and encourage rubber-stamping. Allowing autonomous action on high-consequence, hard-to-reverse cases because routine accuracy is strong creates a different kind of false economy.

Skagway does not currently position this work as broad AI-governance consulting. The Map is relevant only when the automation project exposes a genuine critical-person dependency—when the organization must understand a departing expert’s decision terrain before it can preserve, transfer, or responsibly test the capability. Otherwise, the right partners may be AI assurance, cybersecurity, legal, risk, or domain-specific technical specialists.

Illustrative example

A procurement workflow is authorized to renew contracts below a financial threshold. Historical cases show that a small group of low-value renewals can create a high aggregate dependency when they share one supplier. The evaluation therefore includes two visually similar renewals: one routine, one part of a concentration pattern. The test is whether the system detects the contextual exception and routes it before committing—not whether it can fill the contract fields correctly.

When Skagway is a fit

Skagway Succession is a U.S. executive-succession advisory that captures and transfers the tacit judgment of critical leaders. We are a fit when an organization needs a deliberate, evidence-led process for a critical executive, founder, technical expert, or operator. We are not a replacement for legal, tax, executive-search, compensation, fiduciary, or broad leadership-development advice.

Explore The Map

Glossary

Exception
A case in which context, evidence, consequence, or authority makes the routine response inadequate.
Decision boundary
The condition at which the appropriate interpretation, action, or owner changes.
Evaluation gap
The mismatch that can occur between performance in controlled tests and behavior or utility in real deployment conditions.
Escalation condition
A defined signal or threshold that requires a system to pause, refuse, or transfer the matter to an authorized person.

Sources & further reading

This guide is founder-led analysis. Sources provide background and are not endorsements of Skagway Succession.

Continue the research

What took decades to learn

should not disappear in a day.

The road ahead should remember how the company came this far.