Design the failure path first

The Exception Path Is the Real Workflow

An exception operating model for detecting missing, conflicting, unusual, or failed cases and routing them with evidence to a named owner.

By WhichAI. Published 2026-07-12. Updated 2026-07-12.

Methodology: Editorial synthesis of workflow design patterns and implementation constraints. Public control references provide context, not proof of a deployment or legal advice. Where a versioned evidence pack appears, its evidence class, method, and limitations govern what the artifact can support. Read the full method. Report a correction.

Built for

Operations teams whose clean pilot cases hide ordinary edge work

The decision

Decide how exceptions are detected, classified, enriched, assigned, recovered, and used to improve the system.

Answer first

The happy path proves that a component can work. The exception path determines whether the operation can depend on it. Every failed or ambiguous case needs a reason, evidence, owner, deadline, and recovery action.

WhichAI Solutions diagnostic

Bring this operating problem to the diagnostic

Use WhichAI Solutions when exceptions already consume specialists or managers, failures cross system boundaries, or a missed recovery could affect customers or material operations.

Open the diagnostic

The capacity leak

What the team is doing before anyone calls it a systems problem

Headcount pressure rarely starts with one giant task. It starts when ordinary work is split across inboxes, tabs, handoffs, and undocumented judgment calls. These are the signals to map first.

SIGNAL 01

Exceptions return to a shared inbox with no structured reason or priority.

SIGNAL 02

Staff reconstruct source context and prior attempts before they can resolve a failed case.

SIGNAL 03

Retries create duplicate records or downstream actions because case identity is unstable.

SIGNAL 04

Recurring failure patterns remain hidden in notes and never change preparation or routing rules.

The implementation

The system should prepare the decision, not pretend the decision disappeared

A complete implementation connects the intake, context, transformation, review, and record. The output of one stage becomes the controlled input to the next. A human owns the exceptions and the final consequence.

StageCurrent dragSystem responsibilityHuman responsibilityEvidence kept
1. Exception taxonomyAnything outside the happy path is labeled failed or manual.Define reason codes for missing input, conflicting evidence, unsupported format, policy ambiguity, integration failure, duplicate, and consequential review.Exception owners approve categories, severity, and required response.Taxonomy version, examples, severity, owner, and service target.
2. Detection boundaryFailures are noticed only when a person spots a bad result.Add explicit completeness, conflict, format, permission, duplicate, and downstream-write checks at each stage.Humans validate detection on representative edge cases and identify false alarms.Check result, rule version, test cases, false positives, and missed cases.
3. Exception packetThe resolver receives an error message without business context.Bundle source records, attempted transformations, check results, prior retries, affected action, and suggested next questions.The resolver determines the correct action rather than accepting an inferred fix.Case identifier, source links, attempts, error details, and reviewer decision.
4. Owned recoveryExceptions wait in shared queues or are retried informally.Assign by reason and severity, set a response target, prevent duplicate action, and record resolve, return, reject, or escalate.Named owners approve recovery and consequential downstream action.Assignment, age, action, rationale, downstream reference, and closure time.
5. Failure learningResolved exceptions disappear from view.Review volume and recovery time by cause, then change inputs, rules, training, or scope under version control.Owners approve changes and decide when a case type should leave the automated path.Trend report, root cause, approved change, regression result, and rollback plan.

What the human keeps

The goal is not zero humans. It is zero avoidable preparation around the judgment only a responsible owner should make.

  • Business owners define exception severity, response targets, and consequential escalation.
  • Resolvers interpret ambiguous cases, communicate, and authorize recovery actions.
  • Workflow owners analyze recurring causes and approve controlled changes or scope reduction.

Controls before volume

A workflow is not ready because the happy path worked once. It is ready when access, review, fallback, and evidence are explicit.

  • Give every exception a stable case identifier, reason code, severity, owner, and age.
  • Prevent retries from creating duplicate downstream records or actions.
  • Preserve source evidence and every attempted transformation or recovery.
  • Stop the affected workflow path when a critical control or recovery threshold is breached.

The scorecard

Measure capacity, not activity

A system can produce more messages and still make the operation worse. Measure movement through the workflow, the quality of review, and the load that still reaches a person.

Exception rate by cause

Exceptions per one hundred accepted cases, separated by reason and severity.

Recovery cycle time

Median detection-to-closure time by exception type.

Repeat failure rate

Share of resolved cases that fail again for the same cause or re-enter the queue.

Human recovery load

Resolver minutes per one hundred cases, including context reconstruction and communication.

What a fake implementation looks like here

These patterns create an AI demo while leaving the labor, risk, and accountability in the same place.

  • Using one generic manual queue for every failure type.
  • Retrying failed writes without duplicate protection.
  • Hiding missing or conflicting input behind a plausible default.
  • Optimizing happy-path speed while exception volume consumes more human time.

Two ways to act

Use the path that matches the decision

Questions

What operators ask before they build

What is an exception in an AI workflow?

Any case that cannot safely follow the ordinary preparation and review path because evidence is missing, conflicting, unsupported, consequential, duplicated, or affected by a technical failure.

Should the system suggest a recovery?

It may present approved options and evidence, but the responsible person should choose when the case is ambiguous or consequential. Preserve the selected action and rationale.

When should a case type leave the automated path?

When exception frequency, severity, recovery time, or missed issues exceed the approved threshold and a redesign cannot restore dependable control.

Primary references

Controls should come from the specific operating environment

These are broad public control references, not article-specific evidence, vendor endorsements, or legal advice. Validate the current rules, contracts, system configuration, and organization-specific risk before deployment.

Keep mapping

Related implementation guides

Use this evidence with